Skip to content

Seed2.1: A Next-Generation Agent for Real-World Productivity

Seed2.1 is ByteDance Seed’s successor to Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity — closed-weights API-only family with two tiers, Pro (frontier) and Turbo (lower-latency), targeting “general agents and code engineering” rather than benchmark-chasing. The release page foregrounds an explicit comparison against Claude Opus 4.7, GPT-5.5, and Gemini 3.1 Pro across ~20 benchmarks spanning knowledge, reasoning, agentic high-economic-value work, long-horizon coding, terminal use, debugging, multimodal math/perception/spatial reasoning, and long video understanding (vs Gemini 3.5 Flash / 3.1 Pro for video). Headline placements: SOTA on BabyVision (73.7 vs Gemini 3.1 Pro 54.4), WorldVQA (53.0 vs 44.3), MMLongBench-128K (78.3 vs 70.7), and VideoMME 89.2 / TOMATO 79.5 / Minerva 70.7 / OVOBench 80.7; trails GPT-5.5 on most of the agentic / code / pure-reasoning axes (xDailyBench, ProgramBench, SWE-Atlas, BeyondAIME, SuperGPQA).

  • Two-tier release: Pro (frontier) and Turbo (lower-latency); positioned against Claude Opus 4.7, GPT-5.5, Gemini 3.1 Pro for text/agent benchmarks and against Gemini 3.5 Flash / 3.1 Pro for video [§Evaluation Results].
  • Knowledge — KINA 48.3 (Pro) / 46.6 (Turbo) vs Opus 4.7 46.7, GPT-5.5 52.6, Gemini 3.1 Pro 53.2; SuperGPQA 70.8 / 67.4 vs 68.5 / 72.7 / 76.6 [§Evaluation Results].
  • Reasoning — BeyondAIME 87.0 / 88.0 vs Opus 79.0, GPT-5.5 91.0, Gemini 3.1 Pro 90.0 [§Evaluation Results].
  • High-economic-value agents — Workspace Bench 53.0 / 54.7 (Gemini 3.1 Pro only 32.8, but GPT-5.5 58.7, Opus 55.1); Agent Startup Bench 68.8 / 54.0 (vs Opus 62.3, GPT-5.5 68.1, Gemini 45.7); xDailyBench (white-collar office) 61.0 / 56.4 trails Opus 69.0 and GPT-5.5 73.0 [§Evaluation Results].
  • Long-horizon end-to-end code — NL2Repo-Bench 47.0 / 43.7 trails Opus 58.2, beats GPT-5.5 45.1 and Gemini 33.4; ProgramBench reported as 0/1/50.3 (Pro) and 0/0/49.4 (Turbo) — a triplet format where GPT-5.5 leads at 0.5/5.5/65.9 [§Evaluation Results].
  • Terminal usage — Terminal Bench 2.1 71.0 / 67.6, in line with Opus 71.7, Gemini 3.1 Pro 70.7; trails GPT-5.5 73.8 [§Evaluation Results].
  • Debugging — SWE-Atlas 35.2 / 30.6 trails Opus 38.7 and GPT-5.5 44.7 but beats Gemini 3.1 Pro 23.6 [§Evaluation Results].
  • Multimodal reasoning — MathVision 92.6 / 90.1 (with tool: 94.5 / 92.7) beats Opus 83.1 and Gemini 3.1 Pro 89.2, matches GPT-5.5 92.2; MMMU-Pro 81.6 / 80.1 (w. tool 82.7 / 82.2) leads Opus 74.0, GPT-5.5 81.2, Gemini 3.1 Pro 80.5 [§Evaluation Results].
  • Visual knowledge and perception — WorldVQA 53.0 / 48.6 vs Opus 35.9, GPT-5.5 34.6, Gemini 3.1 Pro 44.3 (large lead); BabyVision 73.7 / 62.9 vs 22.2 / 55.9 / 54.4 (large lead) [§Evaluation Results].
  • Chart QA and spatial reasoning — CharXiv-RQ 85.4 / 82.5 (w. tool 86.4 / 83.6) leads or matches frontier; ERQA spatial reasoning 72.0 / 71.3 leads Opus 52.5, GPT-5.5 64.5, and Gemini 3.1 Pro 70.8 [§Evaluation Results].
  • Long-context multimodal — MMLongBench-128K 78.3 / 76.9 vs Gemini 3.1 Pro 70.7 (Opus and GPT-5.5 not reported) [§Evaluation Results].
  • Video — VideoMME 89.2 / 89 leads Gemini 3.5 Flash 87.2 and Gemini 3.1 Pro 86.7; TOMATO (motion & perception) 79.5 / 56.8 (Pro lead vs Flash 71.9 / Pro 60.4); Minerva (reasoning) 70.7 / 65.9 vs 68.6 / 63.5; OVOBench (streaming) 80.7 / 79.2 vs 64.5 / 64.1 (large lead); VideoSimpleQA 76.4 / 71.4 vs 76 / 70 [§Evaluation Results].
  • ZEROBench — 18.0 / 11.0 (w. tool 22.0 / 20.0) vs Opus 8.0, GPT-5.5 13.0, Gemini 3.1 Pro 12.0 — Seed2.1 leads on this VLM-hard-perception benchmark [§Evaluation Results].
  • Marketing framing positions the model around two consolidated capability stories: general-agent (multi-step office tasks, file processing, tool use) and code engineering (requirement understanding, coding, debugging, validation in real dev workflows) [§Overview, §Showcase].

The release page is a product announcement — no architecture, no training-recipe, no token-count disclosure. What it does specify is the benchmark suite and comparison frontier: 9 text-mode benchmarks (knowledge, reasoning, high-economic-value office work, long-horizon code, terminal, debugging) compared against Claude Opus 4.7, GPT-5.5, Gemini 3.1 Pro; 8 image-mode benchmarks (multimodal math, perception, spatial reasoning, chart QA, long-context multimodal) with the same comparison frontier; and 5 video benchmarks compared against Gemini 3.5 Flash and Gemini 3.1 Pro. The “with tool” parenthetical on MathVision / MMMU-Pro / ZEROBench / CharXiv-RQ indicates a tool-augmented inference path is reported alongside the standalone score — consistent with Seed2.0’s five-axis agentic evaluation framework (Coding / Search / Tool Use / GUI / Deep Research) carried forward from Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity.

The showcase section names two narrow target workflows: (a) generating lesson-plan slides, analyzing complex spreadsheets, and producing industry reports across teaching/office/research; (b) directly generating interactive web pages from floor plans, design mockups, and videos — i.e. multimodal-input → executable code, which is in the spirit of OpenGame: Open Agentic Coding for Games but pitched as a productivity feature rather than a benchmark.

Headline placements (Pro / Turbo) where Seed2.1 leads or matches the frontier:

  • WorldVQA: 53.0 / 48.6 vs Claude Opus 4.7 35.9, GPT-5.5 34.6, Gemini 3.1 Pro 44.3 — ~+9 absolute over the next-best
  • BabyVision: 73.7 / 62.9 vs 22.2 / 55.9 / 54.4 — ~+19 absolute over Gemini 3.1 Pro
  • ERQA (spatial reasoning): 72.0 / 71.3 vs 52.5 / 64.5 / 70.8
  • MMMU-Pro: 81.6 (w. tool 82.7) vs Opus 74.0, GPT-5.5 81.2, Gemini 3.1 Pro 80.5
  • MathVision: 92.6 (w. tool 94.5) vs Opus 83.1, GPT-5.5 92.2, Gemini 3.1 Pro 89.2
  • MMLongBench-128K: 78.3 / 76.9 vs Gemini 3.1 Pro 70.7
  • TOMATO: 79.5 (Pro) vs Gemini 3.5 Flash 71.9, Gemini 3.1 Pro 60.4
  • OVOBench: 80.7 / 79.2 vs 64.5 / 64.1
  • ZEROBench (w. tool): 22.0 / 20.0 vs Opus 8.0, GPT-5.5 13.0, Gemini 3.1 Pro 12.0

Where Seed2.1 trails:

  • xDailyBench (white-collar office): 61.0 / 56.4 vs Opus 69.0, GPT-5.5 73.0
  • NL2Repo-Bench: 47.0 / 43.7 trails Opus 58.2 but beats GPT-5.5 45.1, Gemini 33.4
  • SWE-Atlas: 35.2 / 30.6 trails Opus 38.7, GPT-5.5 44.7
  • BeyondAIME: 87.0 / 88.0 trails GPT-5.5 91.0, Gemini 3.1 Pro 90.0
  • SuperGPQA: 70.8 / 67.4 trails GPT-5.5 72.7, Gemini 3.1 Pro 76.6

The pattern is consistent: Seed2.1 leads on vision-heavy and multimodal-long-context axes (BabyVision, WorldVQA, MMLongBench-128K, OVOBench, TOMATO, ZEROBench, ERQA), holds parity on multimodal reasoning (MathVision, MMMU-Pro, CharXiv), and trails the closed-source frontier on pure-text knowledge, pure-reasoning, and pure-coding axes (xDailyBench, BeyondAIME, SuperGPQA, NL2Repo, SWE-Atlas).

Seed2.1 inherits the candid-gap framing from Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity (Seed2.0 explicitly flagged that it trails Claude on SWE-Evo/NL2Repo and Gemini on SuperGPQA/SimpleQA-Verified) — the same two gaps reappear here: NL2Repo and SWE-Atlas trail Claude Opus 4.7, SuperGPQA trails Gemini 3.1 Pro. Pricing isn’t disclosed on this page, but the Pro/Turbo two-tier replaces Seed2.0’s Pro/Lite/Mini three-tier — the Mini variant is gone or unannounced.

The benchmark composition tells a sharper story than Seed2.0’s: BabyVision, WorldVQA, ERQA, and ZEROBench are all benchmarks the wiki tracks under VLM Perception Failures and adjacent — vision-language benchmarks that expose perception-vs-reasoning gaps in frontier VLMs (Solving Spatial Supersensing Without Spatial Supersensing, BOP-Ask: Object-Interaction Reasoning for Vision-Language Models, Vision Language Models are Biased). Seed2.1’s leads on these (BabyVision +19 over next-best, WorldVQA +9, ERQA +1, ZEROBench +9 with tools) suggest the team is now optimizing against precisely the failure modes the open VLM-perception literature has been characterizing. The TOMATO and OVOBench wins extend the same story into video — TOMATO specifically targets fine-grained motion understanding, which the MotionSight: Boosting Fine-Grained Motion Understanding in Multimodal LLMs team has been calling out as a VLM weak spot.

The Pro/Turbo split itself is worth tracking against the closed-source latency cluster (GPT-5.3-Codex → GPT-5.4 → GPT-5.5 launches all came as paired Thinking/Pro/Instant variants; Gemini’s Flash/Pro/Ultra; Claude’s Sonnet/Opus). Seed2.1’s Turbo is closer to GPT-5.5 / Opus 4.7 on agentic benchmarks than Seed2.0’s Lite was to the comparable frontier, narrowing the cost-vs-quality gap inside the family.

  • Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity — Seed2.0 is the direct predecessor; same candid-gap framing (SWE-class coding trails Claude, long-tail knowledge trails Gemini) carries forward, but pricing and architecture are not disclosed for 2.1
  • Agentic Software Engineering — NL2Repo-Bench / SWE-Atlas / Terminal-Bench 2.1 are the same benchmark spine the open agentic-SWE cluster reports, making Seed2.1 a closed-source datapoint comparable to the Qwen3-Coder-Next / DeepSeek-V3.2 / GLM-4.7 / Step 3.5 Flash Pareto
  • VLM Perception Failures — Seed2.1’s largest leads (BabyVision +19, WorldVQA +9, ZEROBench +9 with tools) are on benchmarks specifically designed to expose VLM perception gaps
  • Unified Multimodal Models — strong video numbers (TOMATO, OVOBench, MMLongBench-128K) suggest deeper multimodal pretraining than typical text-first frontier models
  • Tool-Use Agents — “(w. tool)” parenthetical on MathVision / MMMU-Pro / ZEROBench / CharXiv-RQ indicates tool-augmented inference as a first-class evaluation mode
  • Introducing GPT-5.5 — GPT-5.5 is the closed-source comparison frontier this page picks; a useful cross-reference for which axes Seed2.1 actually catches up on