Skip to content

Introducing GPT-5.5

GPT-5.5 is OpenAI’s April 2026 frontier release, positioned as the successor to GPT-5.4 with the same per-token latency but materially higher capability on agentic-coding, computer-use, knowledge-work, and early-scientific-research workloads. Headline numbers: Terminal-Bench 2.0 82.7%, SWE-Bench Pro 58.6%, Expert-SWE (internal, ~20-hour tasks) 73.1%, OSWorld-Verified 78.7%, GDPval wins/ties 84.9%, Tau2-bench Telecom 98.0%, FrontierMath Tier 4 35.4%, ARC-AGI-2 (Verified) 85.0% — plus a 25–55-point lift over GPT-5.4 on long-context Graphwalks-1M and OpenAI MRCR 512K–1M. The release ships GPT-5.5 and GPT-5.5 Pro to paid ChatGPT and Codex tiers (Codex with 400K context, “Fast mode” trading 2.5× cost for 1.5× decode speed), API “very soon” at 5/5/30 per 1M input/output tokens (Pro at 30/30/180), 1M API context. First OpenAI launch classified High on both bio/chem and cybersecurity under the Preparedness Framework. Notable methodological breadcrumbs (not full disclosures): co-designed on NVIDIA GB200/GB300 NVL72; Codex+GPT-5.5 used internally to rewrite load-balancing/partitioning heuristics in OpenAI’s own inference stack (+20% token-generation throughput); an internal harness over GPT-5.5 found a new asymptotic proof on off-diagonal Ramsey numbers (Lean-verified).

  • GPT-5.5 matches GPT-5.4 per-token latency in real-world serving while reaching materially higher benchmark scores; it also uses “significantly fewer tokens” than GPT-5.4 on Codex tasks at equal or better quality [§“Introducing GPT-5.5”].
  • Terminal-Bench 2.0 82.7% (vs GPT-5.4 75.1%), SWE-Bench Pro 58.6% (vs 57.7%), Expert-SWE 73.1% (vs 68.5%) — all at fewer tokens than GPT-5.4 according to the post body [§“Agentic coding”, Coding table].
  • OSWorld-Verified 78.7% (vs GPT-5.4 75.0%) and GDPval wins-or-ties 84.9% (vs 83.0%, Claude Opus 4.7 82.3%, Gemini 3 Deep Think 82.0%) [§“Knowledge work”, Professional and Computer-use tables].
  • Tau2-bench Telecom 98.0% reported “without prompt tuning” — flagged in the post’s footnote that other labs’ reported numbers used prompt adjustments [§Tool-use table, footnote ***].
  • FrontierMath Tier 1–3 51.7%, Tier 4 35.4%, GPQA Diamond 93.6%, Humanity’s Last Exam (with tools) 52.2%, BixBench 80.5%, GeneBench 25.0% [§Academic table].
  • ARC-AGI-1 (Verified) 95.0%, ARC-AGI-2 (Verified) 85.0% — GPT-5.5 leads GPT-5.4 by ~1.3 and ~11.7 points respectively on the verified tracks [§Abstract reasoning table].
  • Long-context: Graphwalks BFS 1M f1 climbs from 9.4% (GPT-5.4) to 45.4%; OpenAI MRCR v2 8-needle 512K–1M from 36.6% to 74.0% — the largest single-axis gap in the table is at the longest context lengths [§Long context table].
  • Cybersecurity is classified High under the Preparedness Framework; bio/chem also reaches High; GPT-5.5 is the first OpenAI launch with both at High simultaneously [§“Safety, security, and trust”].
  • Co-designed for, trained with, and served on NVIDIA GB200/GB300 NVL72 systems; an internal Codex-driven analysis of weeks of production traffic produced custom partitioning heuristics that “increased token generation speeds by over 20%” — i.e. the model helped improve the infrastructure that serves it [§“Hardware partnership”].
  • An internal version of GPT-5.5 with a custom harness produced a new proof of a longstanding asymptotic fact about off-diagonal Ramsey numbers, later verified in Lean (PDF linked from the post) [§“Scientific and technical research”].
  • Pricing: GPT-5.5 API at 5input/5 input / 30 output per 1M tokens; GPT-5.5 Pro at 30input/30 input / 180 output per 1M tokens; Batch and Flex at half-rate, Priority at 2.5×; 1M-token API context [§“Availability”].
  • Codex availability includes a “Fast mode” that generates tokens 1.5× faster for 2.5× cost — explicit dial trading inference cost for latency, separate from the Pro/Thinking tier [§“Availability”].
  • Capture-the-Flags challenge tasks (internal, “expansion of the hardest CTFs used in system cards”) 88.1% vs GPT-5.4 83.7%; CyberGym 81.8% vs 79.0% [§Cybersecurity table].

The release post discloses no architecture, training-data, or post-training recipe; it is a product/system-card announcement. Methodological breadcrumbs: (a) co-design with NVIDIA GB200/GB300 NVL72 — explicit framing that GPT-5.5 was “trained with and served on” these systems, with “inference as an integrated system, not a set of isolated optimizations” as the engineering thesis; (b) the model assisted in optimizing its own serving infrastructure — Codex analyzed weeks of production traffic and produced custom load-balancing/partitioning heuristic algorithms that replaced static-chunk partitioning, yielding +20% generation throughput, extending the “model that helps build itself” framing from Introducing GPT-5.3-Codex (training-process automation) one further layer into inference-stack automation; (c) deployment-time cybersecurity mitigations include stricter classifiers on cyber-risk requests, a “Trusted Access for Cyber” pilot exposing a cyber-permissive variant (GPT-5.4-Cyber) to verified defenders, and government partnerships for critical-infrastructure defense; (d) “external redteamers” and ~200 trusted early-access partners evaluated the model before release. The reported numbers were run at “reasoning effort xhigh” in a research environment, not production-ChatGPT defaults — so the production numbers may differ.

GPT-5.5 wins every row of the benchmark table over GPT-5.4 with two exceptions: GDPval finance/IB columns where GPT-5.5 vs GPT-5.4 sit at 60.0/56.0 and 88.5/87.3 (clear wins) but Anthropic/Gemini comparison columns reach 61.5 / 88.6 on the same rows, and a couple of MRCR mid-band rows (32K–128K) where GPT-5.4 marginally leads. The largest absolute lifts over GPT-5.4 are at extreme-long context (Graphwalks-1M BFS f1 +36.0, parents-1M +14.1, MRCR 512K–1M +37.4 points), Terminal-Bench 2.0 (+7.6), Tau2 Telecom (+5.2 without prompt tuning), and ARC-AGI-2 Verified (+11.7). On SWE-Bench Pro (the consensus closed/open arena, footnoted with memorization evidence) the lift is only +0.9 over GPT-5.4 and Claude Opus 4.7 at 64.3% still leads — but GPT-5.5 leads on Terminal-Bench 2.0 by 13.3 points over the next-best column (5.4). Cybersecurity capability rises uniformly: +4.4 on the internal CTF expansion, +2.8 on CyberGym. The Ramsey-numbers proof and the GeneBench multi-stage biology-data-analysis numbers (25.0% vs GPT-5.4 19.0%; Gemini 3 Deep Think 33.2% still leads) are the only scientific-research-loop results with a concrete deliverable.

GPT-5.5 is the next closed-frontier datapoint above GPT-5.4 Thinking and GPT-5.4 Pro rolling out in ChatGPT (product announcement) and Introducing GPT-5.3-Codex in the consolidation line the Agentic Software Engineering cluster has been tracking — single unified model, same Pro/Thinking ChatGPT tiering, same SWE-Bench Pro + Terminal-Bench 2.0 + OSWorld + GDPval benchmark quadruple plus added cybersecurity (CTF, CyberGym), academic (FrontierMath, GeneBench, BixBench), and ARC-AGI axes. Two of the framings extend material in already-filed pages: (a) the “model helps build itself” claim has now propagated from training-process automation (GPT-5.3-Codex) to inference-stack optimization (GPT-5.5’s +20% throughput from Codex-authored partitioning heuristics) — the closed-source analog of the closed-loop self-synthesis pattern documented in Close the Loop: Synthesizing Infinite Tool-Use Data via Multi-Agent Role-Playing and Qwen3-Coder-Next Technical Report at the data-generation level; (b) the Ramsey-numbers Lean-verified proof and the GeneBench / BixBench / Bartosz-Naskręcki algebraic-geometry app examples push the inference-time-scaling discussion (Inference-Time Scaling, Recursive Language Models: the paradigm of 2026) toward a “model as research co-scientist” framing with concrete deliverables rather than benchmark wins. The post is also the first OpenAI launch to be classified High on both bio/chem and cybersecurity simultaneously — a Preparedness-Framework precedent worth tracking against the GPT-5.3-Codex High-cyber baseline. Pricing-wise, GPT-5.5 Pro at 30/30/180 per 1M tokens is a strict ~6× markup over GPT-5.5 base for higher-accuracy work, the cleanest filed instance of explicit inference-time-compute pricing tiers.