Skip to content

GLM-5V-Turbo — Vision Coding Model from Z.ai (announcement)

Z.ai announces GLM-5V-Turbo, a vision–coding model positioned as the multimodal sibling of GLM-5. The launch tweet pitches three axes: native multimodal coding (images, videos, design drafts, document layouts as direct inputs), balanced visual + programming capability on multimodal-coding / tool-use / GUI-Agent benchmarks, and deep adaptation for Claude Code and the OpenClaw agent scaffold. No benchmark numbers, system card, or model weights in the tweet — try-it endpoint is chat.z.ai, API docs at docs.z.ai/guides/vlm/glm-5v-turbo, with a Coding Plan trial application form. Filed as a tracker pending an eventual technical report / open-weights release (Haoxiang: “I think they will open-source it later”).

  • The model is named GLM-5V-Turbo and is positioned as a “Vision Coding Model” from Z.ai [tweet body].
  • Natively understands multimodal inputs including images, videos, design drafts, and document layouts [tweet body].
  • Claims “leading performance across core benchmarks for multimodal coding, tool use, and GUI Agents” — no specific benchmarks or numbers given [tweet body].
  • Ships with “deep adaptation” for the Claude Code agent and the OpenClaw agent scaffold [tweet body].
  • Available via the chat.z.ai web product and via API at docs.z.ai/guides/vlm/glm-5v-turbo, with a Coding Plan trial-application Google Form [tweet body].

Not disclosed in the tweet. The naming convention (GLM-5V) suggests a vision-extended variant of the GLM-5 text-only foundation model (744B-total / 40B-active MoE, MLA-256, DSA, trained on 28.5T tokens — see GLM-5: from Vibe Coding to Agentic Engineering), but no architectural confirmation, training-data details, parameter count, or modality-fusion design appears in the tweet. The “Turbo” suffix in Z.ai’s lineup historically denotes a smaller / faster product variant rather than the full flagship, so GLM-5V-Turbo may not be the largest vision model of the GLM-5 generation.

  • No benchmarks are reported in the launch tweet. The “leading performance” claim on multimodal-coding / tool-use / GUI-Agent benchmarks is unquantified.
  • The model is product-accessible at chat.z.ai and via the documented API endpoint; engagement at the time of filing the tweet was ~1.9M views with 252 replies, 932 reposts, 5.7K likes.

This is the multimodal extension of the GLM-5 agentic-coding line (GLM-5: from Vibe Coding to Agentic Engineering) — and lands squarely in the Agentic Software Engineering cluster, but on an axis that no currently-filed agentic-SWE model has formally addressed: multimodal coding with explicit GUI-Agent capability. The closest open peers are Step 3.5 Flash’s edge-cloud collaboration with the Step-GUI on-device agent (Step 3.5 Flash), and the closed-frontier OSWorld / GDPval axis that GPT-5.3-Codex foregrounded (Introducing GPT-5.3-Codex) — both of which the agentic-SWE concept page notes the open cluster has not yet covered directly. The deliberate Claude Code + OpenClaw scaffold compatibility also continues the cross-scaffold tool-call pattern from Qwen3-Coder-Next (Qwen3-Coder-Next Technical Report §4.2.2) and from Step 3.5 Flash’s Claude-Code framing, but extends it to vision-coding workflows. The release is currently API-only; pending an open-weights drop, it does not yet belong on Open foundation-model releases, though Z.ai’s track record with GLM-5 (MIT-licensed, multi-backend serving, seven domestic-GPU adaptation) suggests it likely will.

  • GLM-5: from Vibe Coding to Agentic Engineering — direct text-only sibling; GLM-5V-Turbo is the announced vision-extended variant of this model line.
  • Agentic Software Engineering — extends the cluster to multimodal coding + GUI-Agent capability, an axis no currently-filed open model has covered formally.
  • Step 3.5 Flash — closest open peer on agent-scaffold compatibility (Claude Code) and the only filed agentic-SWE entry that touches GUI agents (Step-GUI edge-cloud orchestration).
  • Introducing GPT-5.3-Codex — closed-frontier reference for the multimodal-coding + computer-use + GUI-agent benchmark quadruple (OSWorld, GDPval) that GLM-5V-Turbo would presumably target.
  • Qwen3-Coder-Next Technical Report — cross-scaffold tool-call training pattern; Z.ai’s “deep adaptation for Claude Code / OpenClaw” is the deployment-side analog.