Skip to content

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity

Seed2.0 is ByteDance’s second-generation closed-weights foundation-model series — three sizes (Pro / Lite / Mini), API-only via Volcano Engine — pitched as a production-deployment-oriented frontier LLM family rather than a benchmark-chasing release. The model card opens by inventorying the surrounding Seed model family (Seed1.6/1.8, Seed1.5-VL, Seed-OSS, Seed-Coder, Seed Diffusion, Seed-Prover) and frames Seed2.0 as the unified general-purpose LLM/agent series targeting “real-world complexity”: LMSYS-Arena human preference (#6 Text, #3 Vision as of 2026-02-16), agentic workflows (coding / search / tool use / GUI / deep research), Olympiad-to-research-grade reasoning (Erdős problems, Scientific Coding), and long-tail professional knowledge. The card is unusually candid about gaps — explicitly naming SWE-Evo and NL2Repo as places Seed2.0 trails Claude, and SuperGPQA / SimpleQA-Verified as places it trails Gemini — and reports headline pricing roughly an order of magnitude below GPT-5.2 / Claude-Opus-4.5 / Gemini-3-Pro at comparable Arena placement.

  • Three-size series: Pro (frontier), Lite (balanced), Mini (high-throughput / low-latency); API-only via Volcano Engine under model id Doubao-Seed-2.0-pro [§1].
  • Pricing per 1M tokens: Pro 0.47prefill/0.47 prefill / 2.37 decode; Lite 0.09/0.09 / 0.53; Mini 0.03/0.03 / 0.31 — vs GPT-5.2 High 1.75/1.75 / 14.00, Claude-Opus-4.5-thinking 5.00/5.00 / 25.00, Gemini-3-Pro 24/2-4 / 12-18 [Table 1].
  • LMSYS Chatbot Arena: #6 on Text Arena overall, #3 on Vision Arena as of 2026-02-16 [§1, footnote 1].
  • 2025 Olympiad math: IMO 2025 35/42 (Gold; threshold ≥35), CMO 2025 114/126 (Gold; threshold ≥87) [Table 2].
  • Capability gaps acknowledged upfront: trails Claude on SWE-Evo and NL2Repo coding; trails Gemini on long-tail-knowledge benchmarks SuperGPQA and SimpleQA-Verified [§1].
  • Evaluation framework for agentic capacity covers five dimensions — Coding Agents, Search Agents, Tool Use, GUI Agents, Deep Research — across benchmarks including SWE-Bench Pro, Terminal-Bench 2.0, SWE-Multilingual, SWE-Evo, τ²-Bench, BFCL-v4, MCP-Mark, VitaBench, BrowseComp, WideSearch, FinSearchComp, DeepConsult, ResearchRubrics, Minedojo-Verified, MM-BrowseComp [§3.3].
  • Vision evaluation spans 50 image benchmarks across 9 categories (MultiModal Math, STEM, Visual Puzzles, Perception & Cognition, General VQA, Pointing & Counting, 2D&3D Spatial, Document & Chart, LongContext) plus 24 video benchmarks across 6 dimensions (knowledge, reasoning, motion & perception, long-video, multi-video, streaming) [§3.2].
  • New Seed-designed benchmarks introduced alongside the release: LPFQA (Long-tail Professional Forum QA), Encyclo-K (book-level professional knowledge with zero/few-shot ICL), HLE-Verified (curated subset of Humanity’s Last Exam after expert review of original-HLE inaccuracies), Ainstain Bench (scientific coding), BABE (interleaved text+visual scientific reasoning, biology), NL2Repo-Bench (end-to-end repo construction from spec) [§3.1.1, §3.4].
  • Advanced evaluation framework organized into four dimensions designed to track real-world agentic workloads: Science Discovery, Vibe Coding, Context Learning, Real-World Tasks [§3.4].
  • MaaS usage in mainland China is concentrated in Internet, Consumer Electronics, Finance, New Retail, Business Services; manufacturing / automotive / communication each below 1% — attributed to capability gaps in prior Seed series rather than market unavailability [§2.1, Fig. 1].
  • Agentic-coding query distribution among Volcano Engine developers: frontend dominates (Vue.js > 3× React adoption in mainland-China developer base), and bug-fixing dominates the task-type split — interpreted as developers using AI for reactive maintenance more than greenfield work [§2.2, Fig. 2].
  • For benchmark fairness, the reported score for each external model is taken as the maximum of (a) the score in its official documentation and (b) ByteDance’s own re-evaluation [§3.3].
  • Multi-format tool-call evaluation is the design target: agentic-evaluation scripts are refactored for execution stability (consolidated Docker images, internal package mirrors, exclusion of non-deterministic / network-dependent / multi-container cases) [§3.3].

The model card is a technical report on the model family and its evaluation methodology rather than a training-recipe paper — no architecture details, no token counts, no RL recipe specifics are disclosed for Seed2.0 itself. What it does specify in detail is the evaluation framework, organized into two layers.

Fundamental capacities (§3.1–§3.3): language (reasoning, complex instruction-following, broad knowledge), vision (50 image + 24 video benchmarks across 9+6 dimensions), and agentic (5 dimensions: Coding Agents, Search Agents, Tool Use, GUI Agents, Deep Research). For each dimension the card enumerates the specific benchmark suite and the comparison frontier models (GPT-5.2 High, Claude-Sonnet-4.5, Claude-Opus-4.5, Gemini-3-Pro High, Gemini-3-Flash High). Two new long-tail-knowledge benchmarks (LPFQA, Encyclo-K) and a curated HLE-Verified subset are introduced to replace what the team argues are noisy or trivia-heavy public benchmarks.

Advanced economically/scientifically valuable tasks (§3.4): a four-axis framework — Scientific Discovery (Ainstain Bench, BABE), Vibe Coding (NL2Repo-Bench targeting full-repo construction from a natural-language spec), Context Learning, Real-World Tasks (in-house benchmarks across Education, Text Classification, Information Extraction). The framing is that the agent era shifts LLMs from “answer isolated prompts” to “drive long-horizon, economically valuable workflows”, and the benchmark suite is constructed to track this shift on concrete failure modes observed in production.

The deployment context (§2) is unusual for a model-card document: usage-distribution statistics by industry and scenario, sourced from the Doubao Collaboration Incentive Program with customer authorization, and an agentic-coding-query distribution by language / framework / development domain / task type derived from Volcano Engine trajectory data. The card frames these as motivation for Seed2.0’s design priorities — visual understanding (because real queries have heavy multimodal content), inference latency (because user satisfaction depends on it), complex-instruction-following (because production tasks chain multi-step constraints), and coding assistance.

  • LMSYS Chatbot Arena (2026-02-16): #6 Text Arena overall, #3 Vision Arena [§1, footnote 1].
  • Math Olympiad (2025): IMO 35/42 Gold, CMO 114/126 Gold [Table 2].
  • Pricing vs frontier (per 1M tokens, USD): Pro 0.47/0.47/2.37, Lite 0.09/0.09/0.53, Mini 0.03/0.03/0.31 vs GPT-5.2 High 1.75/1.75/14.00, Claude-Opus-4.5-thinking 5.00/5.00/25.00, Gemini-3-Pro 24/2-4/12-18, Claude-Sonnet-4.5-thinking 3.00/3.00/15.00, GPT-5.0-mini High 0.25/0.25/2.00, Gemini-3-Flash High 0.501.00/0.50-1.00/3.00 [Table 1]. Roughly an order of magnitude below the frontier closed-source competitors at the Pro tier; two orders of magnitude below at Mini.
  • Acknowledged gaps: trails Claude on SWE-Evo and NL2Repo (coding); trails Gemini on SuperGPQA and SimpleQA-Verified (long-tail knowledge) [§1].

(The truncated fetch did not return the full per-benchmark numbers — third-party summaries report AIME 2025 98.3, HMMT 97.3, Codeforces 3020, SWE-Bench Verified 76.5 for Pro, and Lite within striking distance on WideSearch 74.5 vs 74.7, MMLU-Pro 87.7 vs 87.0, SWE-Bench 73.5 vs 76.5. These are not direct citations from the fetched model card section.)

Seed2.0 is the closed-source mirror image of the open agentic-SWE wave the wiki has been tracking (Agentic Software Engineering) — same benchmark suite (SWE-Bench Pro, Terminal-Bench 2.0, multi-language SWE, τ²-Bench, BFCL-v4), same four-axis production framing (coding + search + tool-use + deep-research), but a tiered closed API at frontier-comparable Arena placement and ~5–10× cheaper than GPT-5.2 / Claude-Opus / Gemini-3-Pro. The model card is also unusually candid about which axes it loses on (SWE-Evo / NL2Repo vs Claude; SuperGPQA / SimpleQA-Verified vs Gemini), which makes it easier to position against than the typical “SOTA everywhere” launch material.

Two specific things worth tracking against the open cluster recipes (Qwen3-Coder-Next Technical Report, Step 3.5 Flash): (1) Seed2.0’s three-size tier with a Mini variant at 0.03/0.03 / 0.31 prefill/decode is the cheapest frontier-API datapoint filed and reinforces the inference-cost-as-headline-axis trend Step 3.5 Flash also flagged; (2) the LPFQA / Encyclo-K / HLE-Verified benchmark designs explicitly target long-tail professional knowledge and book-level mastery — closer to what enterprise MaaS deployments actually do than the math/code competition benchmarks the open cluster tends to lead with. Neither the architecture nor the training recipe is disclosed, so for the open-recipe-comparison angle the card serves as a target rather than a method source.

  • Agentic Software Engineering — Seed2.0’s agentic-coding evaluation suite (SWE-Bench Pro, Terminal-Bench, SWE-Multilingual, SWE-Evo, τ²-Bench, BFCL-v4, MCP-Mark) is a near-perfect superset of what the open agentic-SWE cluster reports.
  • Tool-Use Agents — five-dimension agentic-capacity framework (Coding / Search / Tool Use / GUI / Deep Research) overlaps with this concept’s design surface.
  • Reasoning RL — Olympiad-grade math + Codeforces results suggest a reasoning-RL post-training stage, though the model card discloses no recipe details.
  • Qwen3-Coder-Next Technical Report — Qwen3-Coder-Next is the open-recipe counterpart to Seed2.0 on the SWE-Bench-Pro Pareto; Seed2.0 sets the closed-source target.
  • Introducing GPT-5.3-Codex — GPT-5.3-Codex is the closed-source coding-specialist comparison; Seed2.0 is the closed-source generalist with a closer Lite/Mini tier strategy.
  • Step 3.5 Flash — Step 3.5 Flash also foregrounds inference cost on the headline axis; Seed2.0 Mini extends that framing one tier further down.
  • ByteDance Seed official Seed2.0 page (English): https://seed.bytedance.com/en/seed2
  • Quanquan Gu announcement tweet: https://x.com/QuanquanGu/status/2022560162406707642