Introducing GPT-5.2
GPT-5.2 is OpenAI’s December 11, 2025 frontier release succeeding GPT-5.1 — a three-mode family (Instant, Thinking, Pro) with GPT-5.2 Thinking positioned as the first OpenAI model that performs “at or above a human expert level” on GDPval, OpenAI’s eval of well-specified knowledge work tasks across 44 occupations. Headline pitch: GDPval expert-tie-or-win 70.9%, plus a state-of-the-art long-context reasoning claim, “major improvements” in spreadsheet creation/analysis/formatting, and early gains in slideshow creation. Rolls out to Plus/Pro/Business/Enterprise on launch day with Free/Go a day later; API and Codex available immediately for all developers. Pro is OpenAI’s “smartest and most trustworthy” tier with extended reasoning for difficult questions; a separately-trained GPT-5.2-Codex variant followed. This is the closed-frontier datapoint immediately below Introducing GPT-5.5 on the GPT-5 release timeline (GPT-5.4 Thinking and GPT-5.4 Pro rolling out in ChatGPT (product announcement) (GPT-5.4) and Introducing GPT-5.3-Codex sit between them) and the first GPT-5-family launch to disclose GDPval at the headline level rather than as a side benchmark.
Key claims
Section titled “Key claims”- GPT-5.2 Thinking is OpenAI’s first model to perform at human-expert level on GDPval, beating or tying top industry professionals on 70.9% of comparisons across 44 occupations, with tasks including presentations, spreadsheets, and other artifacts [§“GPT‑5.2 Thinking is the best model yet for real-world, professional use”; tweet].
- GPT-5.2 brings significant improvements in general intelligence, long-context understanding, agentic tool-calling, and vision, with GPT-5.2 Thinking framed as state-of-the-art on long-context reasoning [§“Overall”].
- GPT-5.2 Thinking’s headline domain gains are spreadsheet creation/analysis/formatting (major improvements), slideshow creation (early gains), coding, summarizing long documents, answering questions about uploaded files, and step-by-step math/logic [§“GPT‑5.2 Thinking”].
- GPT-5.2 Instant’s positioning shifts toward warmer everyday work and learning while keeping GPT-5.1’s conversational tone — clearer up-front explanations, better how-tos and walk-throughs, stronger technical writing/translation, study and career-guidance support [tweet].
- GPT-5.2 Pro is the “smartest and most trustworthy” tier for difficult questions, with early-testing reports of fewer major errors and stronger programming performance, and is OpenAI’s “best model for assisting and accelerating scientists” [§“GPT‑5.2 Pro”; tweet].
- Rollout: GPT-5.2 Instant/Thinking/Pro begin rolling out to Plus, Pro, Business, and Enterprise on launch day; Free and Go users one day later; API access available to all developers at launch [§“Rollout”; tweet].
- A Codex-optimized GPT-5.2 variant is announced as upcoming “in the coming weeks” — GPT-5.2 is expected to work well in Codex out of the box but a specialized variant is the deployment target for that surface [§“Rollout”].
- Safety: builds on GPT-5-era “safe completions” research, with targeted improvements in sensitive-conversation handling (signs of suicide/self-harm, mental health distress, emotional reliance) and fewer undesirable responses than GPT-5.1 / GPT-5 Instant + Thinking on those evals [§“Safety”].
- Age-prediction-driven content protections for under-18 users are in early rollout, layered on top of existing under-18 handling and parental controls [§“Safety”].
- No current plans to deprecate GPT-5.1, GPT-5, or GPT-4.1 in the API, with advance notice promised for any future deprecation [§“Rollout”].
- Reported benchmark numbers were run at “maximum available reasoning effort” (xhigh for GPT-5.2 Thinking and Pro, high for GPT-5.1 Thinking), except for the professional evals — i.e. the headline numbers are not production-ChatGPT-default settings [§“Benchmarks”].
- GPT-5.1 in ChatGPT remains available to paid users as a legacy model for three months [tweet].
Method
Section titled “Method”The release post is a product/system-card announcement with no architecture, training-data, or post-training-recipe disclosure. The disclosed framing is incremental: GPT-5.2 is positioned as “one step in an ongoing series of improvements” rather than a base-model overhaul, building on GPT-5’s “safe completions” research and continuing the Instant / Thinking / Pro tiering. The only methodological breadcrumbs are (a) the existence of a Codex-specific variant in the pipeline, (b) targeted post-training work on sensitive-conversation responses, and (c) an explicit “maximum reasoning effort” caveat on the reported benchmark numbers (xhigh for Thinking/Pro, high for the GPT-5.1 Thinking comparison) — meaning the GDPval and other headline figures are not what production ChatGPT delivers by default. Subsequent OpenAI posts (Introducing GPT-5.5) confirm that GPT-5.2 reuses the GPT-5-family router architecture across Instant/Thinking/Pro variants with reasoning-effort levels (xhigh through non-reasoning) as the inference-compute dial.
Results
Section titled “Results”The headline result is GDPval expert-tie-or-win at 70.9% for GPT-5.2 Thinking — the first OpenAI model to cross the “at or above human expert” line on the eval. The tweet thread frames this as the central pitch. Beyond GDPval the release post claims state-of-the-art on long-context reasoning and major improvements in spreadsheet creation/analysis/formatting plus early gains in slideshow creation, without numbers in the public post body for those axes. GPT-5.2 Pro is positioned for programming and assisting scientists with “fewer major errors” from early testers, again without disclosed benchmark numbers in the public post body. Subsequent OpenAI launches close the gap: Introducing GPT-5.5 retroactively publishes a full GPT-5.4 vs. GPT-5.5 benchmark grid (Terminal-Bench 2.0, SWE-Bench Pro, Expert-SWE, OSWorld-Verified, GDPval, FrontierMath, ARC-AGI, Graphwalks, MRCR) but does not back-fill GPT-5.2 numbers, leaving GDPval as the only quantified headline for this release. Wikipedia and contemporaneous reporting flag GPT-5.2’s release as occurring ~3 weeks after Gemini 3 Pro and as accelerated by an internal “Code Red” memo prompted by Gemini’s then-leading multimodal numbers.
Why it’s interesting
Section titled “Why it’s interesting”GPT-5.2 is the closed-frontier reference point Introducing GPT-5.5 uses as its longest-running comparison baseline (every Long-context, Graphwalks, MRCR, and FrontierMath row in the GPT-5.5 grid is anchored against GPT-5.4 → GPT-5.2 → GPT-5.1), so filing it makes the closed-source GPT-5-family lineage on the Agentic Software Engineering page complete: GPT-5.2 (Dec 2025) → GPT-5.3-Codex (Introducing GPT-5.3-Codex, Feb 2026) → GPT-5.4 (GPT-5.4 Thinking and GPT-5.4 Pro rolling out in ChatGPT (product announcement), Mar 2026) → GPT-5.5 (Introducing GPT-5.5, Apr 2026). Two framings in this release become load-bearing later: (a) GDPval as the headline “professional knowledge work” axis — GPT-5.2 introduces it at 70.9% expert-tie-or-win, GPT-5.3-Codex picks it up as one corner of the agentic quadruple (SWE-Bench Pro + Terminal-Bench 2.0 + OSWorld + GDPval), and GPT-5.5 reports a full 44-occupation column at 84.9%; (b) the Instant/Thinking/Pro tiering with reasoning-effort as an explicit inference-compute dial — first deployed at the product surface here, then continued through the next three GPT-5-family releases and connecting directly to the Inference-Time Scaling cluster’s “Pro vs base” sequence-budget axis. GPT-5.2’s “first OpenAI model at human-expert level on GDPval” framing also fits the Reasoning RL page’s broader trend toward reasoning-model post-training pulling closed-source models past professional baselines on knowledge work, in parallel to the open MoE Pareto frontier (Step 3.5 Flash, Qwen3-Coder-Next Technical Report) which still trails on the OSWorld/GDPval corners of the quadruple.
See also
Section titled “See also”- Introducing GPT-5.5 — the next-but-two GPT-5-family release; back-references GPT-5.2 as the long-context / FrontierMath / Graphwalks baseline across its full benchmark grid
- GPT-5.4 Thinking and GPT-5.4 Pro rolling out in ChatGPT (product announcement) — GPT-5.4 launch tweet; the intermediate consolidation step between GPT-5.2 and GPT-5.5
- Introducing GPT-5.3-Codex — GPT-5.3-Codex (the announced Codex-specific variant of the GPT-5.2 generation, released ~2 months later); first introduces the SWE-Bench Pro + Terminal-Bench 2.0 + OSWorld + GDPval quadruple that this release seeds
- GPT-5.3 Instant: Smoother, more useful everyday conversations — the Instant-tier follow-on tracing the Instant variant’s evolution from GPT-5.2 onward
- Agentic Software Engineering — closed-frontier lineage GPT-5.2 anchors; GDPval at 70.9% is the first filed quantified datapoint on the professional-knowledge-work axis of the agentic quadruple
- Inference-Time Scaling — Pro vs Instant vs Thinking with reasoning-effort as the explicit compute dial; GPT-5.2 is the first GPT-5-family release where the dial is the product packaging
- Reasoning RL — “first OpenAI model at human-expert level on GDPval” sits squarely on this concept’s frontier-models-on-knowledge-work axis
- Open foundation-model releases — closed-but-API-accessible cohort that this concept page’s open releases compare against; GPT-5.2 is what Kimi K2.5’s benchmark table directly benchmarks against