Skip to content

Build with Gemini 3 Flash, frontier intelligence that scales with you

Google ships Gemini 3 Flash as a reasoning + agentic model positioned at ~1/4 the cost of Gemini 3 Pro and ~3× faster than 2.5 Pro, while outperforming 2.5 Pro on benchmarks. Headline numbers: 90.4% GPQA Diamond, 33.7% Humanity’s Last Exam (no tools), 78% SWE-Bench Verified — the SWE-Bench number is reported to exceed Gemini 3 Pro’s own agentic coding score. Priced at 0.50/Minput,0.50/M input, 3/M output, with context caching (claimed 90% reductions for repeated tokens) and Batch API (50% off). Rolling out via Gemini API, AI Studio, Antigravity, Gemini CLI, Android Studio, and Vertex AI.

  • Gemini 3 Flash hits 90.4% on GPQA Diamond and 33.7% on Humanity’s Last Exam without tools [§Smarter, faster and ready for production at scale].
  • Gemini 3 Flash scores 78% on SWE-Bench Verified, which the post claims outperforms Gemini 3 Pro’s own agentic coding score while running faster [§For coding].
  • Priced at 0.50/Minputtokensand0.50/M input tokens and 3/M output tokens; audio input remains $1/M [§Smarter, faster and ready for production at scale].
  • Standard context caching reduces cost by up to 90% on repeated-token thresholds; Batch API offers 50% cost savings on asynchronous workloads [§Smarter, faster and ready for production at scale].
  • Even at the lowest “thinking level,” 3 Flash “often outperforms previous versions with high thinking levels” — a thinking-budget claim measured against earlier Gemini Flash generations [§Smarter, faster and ready for production at scale].
  • 3 Flash is positioned as ~3× faster than 2.5 Pro on the Artificial Analysis benchmark while outperforming it [§Smarter, faster and ready for production at scale].
  • Available immediately in Google AI Studio, Gemini API, Antigravity, Gemini CLI, Android Studio, Gemini Code Assist, and Vertex AI [§Get started with Gemini 3 Flash].
  • Resemble AI case study claims 4× faster multimodal analysis on deepfake forensic data vs. Gemini 2.5 Pro [§For deepfake detection].

Not disclosed. This is a product-launch post — no architecture, training recipe, or scaling-law information is given. The post is structured around (a) headline benchmark numbers vs. prior Gemini generations, (b) the Pareto-frontier framing of cost × speed × quality, (c) integration surface (which Google products the model ships into), and (d) early-customer case studies (Astrocade for game generation, Latitude for character agents, Resemble for deepfake forensics, Harvey for legal document analysis).

The “thinking level” knob is referenced as a developer-facing control (ai.google.dev/gemini-api/docs/thinking#thinking-levels) but not characterized — it appears to expose the model’s reasoning-token budget. The post also flags a new Interactions API and the requirement to “circulate thoughts in the API” via thought signatures, suggesting Gemini 3 Flash expects multi-turn agentic loops to thread reasoning state explicitly.

Three reported benchmark numbers:

  • GPQA Diamond: 90.4%.
  • Humanity’s Last Exam (no tools): 33.7%.
  • SWE-Bench Verified (agentic coding): 78%.

Pricing and serving numbers:

  • 0.50/0.50 / 3.00 per 1M input / output tokens; $1/M audio input.
  • Up to 90% cost reduction via standard context caching at repeated-token thresholds.
  • 50% cost reduction via Batch API for async workloads.
  • ~3× speedup over Gemini 2.5 Pro (Artificial Analysis benchmark).

The post does not include head-to-head numbers against GPT-5.x, Claude Opus 4.x, DeepSeek V3.2, or Kimi K2.5 (the cohort the open-foundation-releases page tracks Gemini 3 Pro against). The 78% SWE-Bench Verified > Gemini 3 Pro claim is the most interesting line on the page — a Flash-tier model beating the same family’s Pro tier on an agentic benchmark is the kind of result that earlier “Flash” releases did not produce.

This is the Flash-tier shoe to drop after Nano Banana 2 / Gemini 3.1 Flash Image (Google product announcement) (Nano Banana 2 / Gemini 3.1 Flash Image) and Gemini 3.1 Flash Live: realtime voice and vision agent model (Google product announcement) (Gemini 3.1 Flash Live) — the same model family, now with the reasoning + agentic text checkpoint shipped openly. Together they fill in what Sundar Pichai 🤔🤔 reply teasing Gemini 3 release timing teased and what Gemini 3 Deep Think: Advancing science, research and engineering announced on the Pro side.

The benchmark line that matters for Luma research is the 78% SWE-Bench Verified at Flash-tier pricing/latency: this is the same agentic-coding axis where the Open foundation-model releases concept page tracks Kimi K2.5 reporting frontier parity (60.6–78.4 BrowseComp) against closed Gemini 3 Pro. A Flash-tier Gemini now beating its own Pro tier on SWE-Bench complicates that comparison — it suggests the “frontier coding model” target is moving inside Google’s own lineup faster than the open-source benchmark cycle.

The thinking-level + thought-signatures plumbing is also worth noting: it’s the most explicit productization yet of Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities-style adjustable reasoning budgets, exposed as a developer knob alongside an Interactions API for multi-turn agent state.