GLM-5.1 — Next Level of Open Source (Z.ai announcement)
Z.ai announces GLM-5.1, an incremental open-source release on top of GLM-5 (GLM-5: from Vibe Coding to Agentic Engineering) targeting long-horizon agentic coding. Headline tweet claims: #1 in open source and #3 globally on SWE-Bench Pro + Terminal-Bench + NL2Repo, SWE-Bench Pro = 58.4, and autonomous 8-hour runs with self-review refinement across thousands of iterations. Two demo scenarios are pitched: building a Linux Desktop from scratch (8-hour self-review loop), and a Vector-DB-Bench optimization run reaching 21.5k QPS over 600+ iterations and 6,000+ tool calls (6× a standard 50-turn session). No system card, weights link, or training-recipe details in the launch tweet — filed as a tracker pending a technical report.
Key claims
Section titled “Key claims”- GLM-5.1 reports #1 in open source and #3 globally across SWE-Bench Pro, Terminal-Bench, and NL2Repo [tweet body].
- Reports 58.4 on SWE-Bench Pro, framed as a “significant leap in coding and agentic performance” over GLM-5 [tweet body].
- Designed for long-horizon tasks: “Runs autonomously for 8 hours, refining strategies through thousands of iterations” [tweet body].
- Linux Desktop demo: 8-hour self-review loop autonomously refining features, styling, and interactions to build a functional desktop environment [tweet body].
- Vector-DB-Bench demo: reached 21.5k QPS over 600+ iterations and 6,000+ tool calls — claimed at 6× the performance of a standard 50-turn session [tweet body].
- The thread also (separately) announces GLM-OCR (0.9B SOTA OCR with CogViT visual encoder + GLM-0.5B decoder, two-stage layout+parallel pipeline based on PP-DocLayout-V3), GLM-4.6V (106B vision-language flagship with 128K context and first-ever native Function Calling in the GLM vision family), and GLM-4.6V-Flash (9B ultra-fast variant) [tweet thread body].
Method
Section titled “Method”Not disclosed in the launch tweet. Naming convention places GLM-5.1 as a post-training refresh on top of GLM-5’s 744B-total / 40B-active MoE backbone with DeepSeek Sparse Attention and the sequential Reasoning RL → Agentic RL → General RL pipeline (GLM-5: from Vibe Coding to Agentic Engineering), but no architectural or training-recipe details are confirmed. The “8 hours autonomous, thousands of iterations” framing implies further RL training against long-horizon, self-reviewing environments — likely an extension of GLM-5’s asynchronous slime-based agent-RL infrastructure — but no algorithm name or environment detail appears in the tweet.
Results
Section titled “Results”- SWE-Bench Pro: 58.4 (tweet headline figure). Context: prior open-cluster leaders on SWE-Bench Pro include Qwen3-Coder-Next 44.3, DeepSeek-V3.2 40.9, GLM-4.7 40.6 (Qwen3-Coder-Next Technical Report Fig. 1); closed-source GPT-5.3-Codex sets the upper bound at 56.8 (Introducing GPT-5.3-Codex). If reproducible, GLM-5.1’s 58.4 would be the first open-weights number to cross 50 on SWE-Bench Pro and to exceed the GPT-5.3-Codex baseline.
- Terminal-Bench, NL2Repo: ranking-only (#1 open / #3 global); no per-benchmark numerical figures in the tweet.
- Vector-DB-Bench demo (internal): 21.5k QPS, 6× a 50-turn session, 600+ iterations / 6,000+ tool calls.
- Linux Desktop demo (internal): 8-hour autonomous build of a functional desktop environment.
- Engagement at filing: tweet thread visible at chat.z.ai with linked share URL
chat.z.ai/s/f4389582-bcd…. No reproducible eval harness or system card linked from the tweet.
Why it’s interesting
Section titled “Why it’s interesting”GLM-5.1 is the first filed open-weights model to claim a SWE-Bench Pro number above the GPT-5.3-Codex closed-source bar (58.4 vs 56.8) — if the number holds up under independent reproduction, it inverts the closed-vs-open ordering that Agentic Software Engineering currently tracks (where the open Pareto frontier sits at Qwen3-Coder-Next 44.3 and GPT-5.3-Codex sets the upper bound). The 8-hour autonomous self-review loop with thousands of iterations is a sharper claim on the “long-horizon agent” axis than any currently-filed open model: Vending-Bench 2 (1-year simulated business) for GLM-5: from Vibe Coding to Agentic Engineering is the only comparable long-horizon datapoint, and that was a single-task simulation; the Linux Desktop demo and Vector-DB-Bench 6,000+ tool-call rollout are concrete agentic-trajectory scales no other open release has highlighted. The tweet thread also continues the multi-product simultaneous-release packaging pattern from Open foundation-model releases — GLM-OCR (0.9B perception model), GLM-4.6V (106B flagship VLM with native Function Calling), and GLM-4.6V-Flash (9B) appearing in the same thread as the flagship GLM-5.1 launch. This is structurally similar to NVIDIA’s Jan 2026 multi-domain bundle but at single-org scope. Concrete weights and a technical report would resolve whether the 58.4 number is methodologically comparable to the other entries on this Pareto.
See also
Section titled “See also”- GLM-5: from Vibe Coding to Agentic Engineering — direct predecessor (GLM-5: 744B-total / 40B-active MoE with DSA + sequential Reasoning RL → Agentic RL → General RL); GLM-5.1 is the announced next iteration.
- GLM-5V-Turbo — Vision Coding Model from Z.ai (announcement) — sibling Z.ai announcement (GLM-5V-Turbo vision-coding model); the GLM-5.1 thread also announces GLM-4.6V / GLM-4.6V-Flash, extending the GLM vision-coding line that GLM-5V-Turbo opened.
- Agentic Software Engineering — claims #1 open / #3 global on the SWE-Bench Pro + Terminal-Bench + NL2Repo triple; the 58.4 SWE-Bench Pro figure, if reproducible, sits above GPT-5.3-Codex’s 56.8 closed-source upper bound currently tracked on that page.
- Qwen3-Coder-Next Technical Report — closest open-cluster peer on SWE-Bench Pro Pareto (44.3 vs the claimed 58.4).
- Introducing GPT-5.3-Codex — closed-frontier reference (56.8 SWE-Bench Pro); GLM-5.1’s headline figure crosses this bar.
- Step 3.5 Flash — open-MoE peer also reporting SWE-bench Verified 74.4 and Terminal-Bench 2.0 51.0; shares the long-horizon agent-RL framing.
- Open foundation-model releases — the GLM-5.1 thread bundles GLM-OCR / GLM-4.6V / GLM-4.6V-Flash in a single announcement, extending the multi-product simultaneous-release pattern.