Gemini 3 Deep Think: Advancing science, research and engineering
Google blog post announcing a major upgrade to Gemini 3 Deep Think, a specialized reasoning mode (closed model, closed methodology). The release sets a new top score on Humanity’s Last Exam without tools (48.4%), claims ARC-AGI-2 at 84.6% verified by the ARC Prize Foundation, a Codeforces Elo of 3455, and gold-medal performance on IMO 2025 / IPhO 2025 / IChO 2025 written sections plus 50.5% on CMT-Benchmark. Deep Think is available in the Gemini app to Google AI Ultra subscribers, and for the first time via the Gemini API under an early-access program. No architectural, training, or inference-time-compute details are disclosed.
Key claims
Section titled “Key claims”- A new top score of 48.4% on Humanity’s Last Exam without tools [§Elevating reasoning with mathematical and algorithmic rigor].
- 84.6% on ARC-AGI-2, verified by the ARC Prize Foundation [§Elevating reasoning with mathematical and algorithmic rigor].
- A Codeforces Elo of 3455 [§Elevating reasoning with mathematical and algorithmic rigor].
- Gold-medal level performance on the IMO 2025, and on the written sections of the 2025 International Physics Olympiad and Chemistry Olympiad [§Elevating reasoning with mathematical and algorithmic rigor, §Navigating complex scientific domains].
- 50.5% on CMT-Benchmark (advanced theoretical physics) [§Navigating complex scientific domains].
- Deep Think is now shipped via the Gemini API to select researchers, engineers and enterprises under an early-access program — the first API exposure for the Deep Think mode [§Early Access Program availability].
Method
Section titled “Method”Not disclosed. The post is a product announcement; it does not describe the architecture, training procedure, or how Deep Think allocates extra inference-time compute. The framing is that Deep Think is a “specialized reasoning mode” that goes beyond the base Gemini 3, in line with prior Deep Think variants that the post cites as having achieved gold-medal standards at IMO and ICPC and as enabling research-level mathematics-exploration agents. The post emphasizes broadening from competition math/coding into chemistry, physics, materials science, and physical-component design, with anecdotes from external testers (Rutgers / Duke Wang Lab / a Google R&D lead) rather than benchmark deltas on the broader scientific tasks.
Results
Section titled “Results”Headline benchmark numbers on a single closed model:
- Humanity’s Last Exam: 48.4% without tools (claimed new top score).
- ARC-AGI-2: 84.6% (claimed unprecedented; verified by the ARC Prize Foundation).
- Codeforces: 3455 Elo.
- IMO 2025: gold-medal level.
- IPhO 2025 (written): gold-medal level.
- IChO 2025 (written): gold-medal level.
- CMT-Benchmark: 50.5%.
No comparison baselines are reported in the post (e.g. base Gemini 3 without Deep Think, prior Deep Think versions, or competing frontier models). Application anecdotes — identifying a logical flaw in a peer-reviewed mathematics paper, designing a recipe for crystal thin films > 100 μm, sketch-to-3D-printable-component synthesis — are qualitative.
Why it’s interesting
Section titled “Why it’s interesting”A closed-model datapoint on the frontier of test-time-compute reasoning, useful primarily as a comparison baseline rather than a methods source. Two specific connections to the wiki: (a) the K2.5 release benchmark table (Kimi K2.5: Visual Agentic Intelligence) names Gemini 3 Pro as one of its five closed-frontier comparisons — Deep Think is the reasoning-mode sibling that K2.5 implicitly competes with on browse/math benchmarks; (b) Deep Think is by construction an inference-time-scaling product, but the post discloses none of the six axes catalogued under Inference-Time Scaling (sequential CoT, scaffold/REPL, interaction depth, KV-cache compression, learned parallel orchestration, generation-side iterative refinement, test-time training of weights) — so this is the closed-product datapoint that the open-source axes are being measured against, not a contribution to the taxonomy itself. The ARC-AGI-2 jump (84.6%, verified) is the most independently auditable claim and the one worth tracking against subsequent open-source attempts.
See also
Section titled “See also”- Inference-Time Scaling — Deep Think is a frontier closed product in this space, but discloses no mechanism
- Reasoning RL — implied training paradigm for reasoning-specialized variants; not described in the post
- Open foundation-model releases — counter-point: this is a closed, gated, API-only release with neither weights nor methodology
- Kimi K2.5: Visual Agentic Intelligence — open frontier release that explicitly benchmarks against Gemini 3 Pro on shared tasks
- DeepSeek-OCR 2: Visual Causal Flow — open release that beats Gemini-3 Pro at matched token budget on OmniDocBench v1.5; the converse comparison on closed-side reasoning benchmarks