Skip to content

Gemini 3 Deep Think: Advancing science, research and engineering

Google blog post announcing a major upgrade to Gemini 3 Deep Think, a specialized reasoning mode (closed model, closed methodology). The release sets a new top score on Humanity’s Last Exam without tools (48.4%), claims ARC-AGI-2 at 84.6% verified by the ARC Prize Foundation, a Codeforces Elo of 3455, and gold-medal performance on IMO 2025 / IPhO 2025 / IChO 2025 written sections plus 50.5% on CMT-Benchmark. Deep Think is available in the Gemini app to Google AI Ultra subscribers, and for the first time via the Gemini API under an early-access program. No architectural, training, or inference-time-compute details are disclosed.

  • A new top score of 48.4% on Humanity’s Last Exam without tools [§Elevating reasoning with mathematical and algorithmic rigor].
  • 84.6% on ARC-AGI-2, verified by the ARC Prize Foundation [§Elevating reasoning with mathematical and algorithmic rigor].
  • A Codeforces Elo of 3455 [§Elevating reasoning with mathematical and algorithmic rigor].
  • Gold-medal level performance on the IMO 2025, and on the written sections of the 2025 International Physics Olympiad and Chemistry Olympiad [§Elevating reasoning with mathematical and algorithmic rigor, §Navigating complex scientific domains].
  • 50.5% on CMT-Benchmark (advanced theoretical physics) [§Navigating complex scientific domains].
  • Deep Think is now shipped via the Gemini API to select researchers, engineers and enterprises under an early-access program — the first API exposure for the Deep Think mode [§Early Access Program availability].

Not disclosed. The post is a product announcement; it does not describe the architecture, training procedure, or how Deep Think allocates extra inference-time compute. The framing is that Deep Think is a “specialized reasoning mode” that goes beyond the base Gemini 3, in line with prior Deep Think variants that the post cites as having achieved gold-medal standards at IMO and ICPC and as enabling research-level mathematics-exploration agents. The post emphasizes broadening from competition math/coding into chemistry, physics, materials science, and physical-component design, with anecdotes from external testers (Rutgers / Duke Wang Lab / a Google R&D lead) rather than benchmark deltas on the broader scientific tasks.

Headline benchmark numbers on a single closed model:

  • Humanity’s Last Exam: 48.4% without tools (claimed new top score).
  • ARC-AGI-2: 84.6% (claimed unprecedented; verified by the ARC Prize Foundation).
  • Codeforces: 3455 Elo.
  • IMO 2025: gold-medal level.
  • IPhO 2025 (written): gold-medal level.
  • IChO 2025 (written): gold-medal level.
  • CMT-Benchmark: 50.5%.

No comparison baselines are reported in the post (e.g. base Gemini 3 without Deep Think, prior Deep Think versions, or competing frontier models). Application anecdotes — identifying a logical flaw in a peer-reviewed mathematics paper, designing a recipe for crystal thin films > 100 μm, sketch-to-3D-printable-component synthesis — are qualitative.

A closed-model datapoint on the frontier of test-time-compute reasoning, useful primarily as a comparison baseline rather than a methods source. Two specific connections to the wiki: (a) the K2.5 release benchmark table (Kimi K2.5: Visual Agentic Intelligence) names Gemini 3 Pro as one of its five closed-frontier comparisons — Deep Think is the reasoning-mode sibling that K2.5 implicitly competes with on browse/math benchmarks; (b) Deep Think is by construction an inference-time-scaling product, but the post discloses none of the six axes catalogued under Inference-Time Scaling (sequential CoT, scaffold/REPL, interaction depth, KV-cache compression, learned parallel orchestration, generation-side iterative refinement, test-time training of weights) — so this is the closed-product datapoint that the open-source axes are being measured against, not a contribution to the taxonomy itself. The ARC-AGI-2 jump (84.6%, verified) is the most independently auditable claim and the one worth tracking against subsequent open-source attempts.