Build with Gemini 3 Flash, frontier intelligence that scales with you
Google ships Gemini 3 Flash as a reasoning + agentic model positioned at ~1/4 the cost of Gemini 3 Pro and ~3× faster than 2.5 Pro, while outperforming 2.5 Pro on benchmarks. Headline numbers: 90.4% GPQA Diamond, 33.7% Humanity’s Last Exam (no tools), 78% SWE-Bench Verified — the SWE-Bench number is reported to exceed Gemini 3 Pro’s own agentic coding score. Priced at 3/M output, with context caching (claimed 90% reductions for repeated tokens) and Batch API (50% off). Rolling out via Gemini API, AI Studio, Antigravity, Gemini CLI, Android Studio, and Vertex AI.
Key claims
Section titled “Key claims”- Gemini 3 Flash hits 90.4% on GPQA Diamond and 33.7% on Humanity’s Last Exam without tools [§Smarter, faster and ready for production at scale].
- Gemini 3 Flash scores 78% on SWE-Bench Verified, which the post claims outperforms Gemini 3 Pro’s own agentic coding score while running faster [§For coding].
- Priced at 3/M output tokens; audio input remains $1/M [§Smarter, faster and ready for production at scale].
- Standard context caching reduces cost by up to 90% on repeated-token thresholds; Batch API offers 50% cost savings on asynchronous workloads [§Smarter, faster and ready for production at scale].
- Even at the lowest “thinking level,” 3 Flash “often outperforms previous versions with high thinking levels” — a thinking-budget claim measured against earlier Gemini Flash generations [§Smarter, faster and ready for production at scale].
- 3 Flash is positioned as ~3× faster than 2.5 Pro on the Artificial Analysis benchmark while outperforming it [§Smarter, faster and ready for production at scale].
- Available immediately in Google AI Studio, Gemini API, Antigravity, Gemini CLI, Android Studio, Gemini Code Assist, and Vertex AI [§Get started with Gemini 3 Flash].
- Resemble AI case study claims 4× faster multimodal analysis on deepfake forensic data vs. Gemini 2.5 Pro [§For deepfake detection].
Method
Section titled “Method”Not disclosed. This is a product-launch post — no architecture, training recipe, or scaling-law information is given. The post is structured around (a) headline benchmark numbers vs. prior Gemini generations, (b) the Pareto-frontier framing of cost × speed × quality, (c) integration surface (which Google products the model ships into), and (d) early-customer case studies (Astrocade for game generation, Latitude for character agents, Resemble for deepfake forensics, Harvey for legal document analysis).
The “thinking level” knob is referenced as a developer-facing control (ai.google.dev/gemini-api/docs/thinking#thinking-levels) but not characterized — it appears to expose the model’s reasoning-token budget. The post also flags a new Interactions API and the requirement to “circulate thoughts in the API” via thought signatures, suggesting Gemini 3 Flash expects multi-turn agentic loops to thread reasoning state explicitly.
Results
Section titled “Results”Three reported benchmark numbers:
- GPQA Diamond: 90.4%.
- Humanity’s Last Exam (no tools): 33.7%.
- SWE-Bench Verified (agentic coding): 78%.
Pricing and serving numbers:
- 3.00 per 1M input / output tokens; $1/M audio input.
- Up to 90% cost reduction via standard context caching at repeated-token thresholds.
- 50% cost reduction via Batch API for async workloads.
- ~3× speedup over Gemini 2.5 Pro (Artificial Analysis benchmark).
The post does not include head-to-head numbers against GPT-5.x, Claude Opus 4.x, DeepSeek V3.2, or Kimi K2.5 (the cohort the open-foundation-releases page tracks Gemini 3 Pro against). The 78% SWE-Bench Verified > Gemini 3 Pro claim is the most interesting line on the page — a Flash-tier model beating the same family’s Pro tier on an agentic benchmark is the kind of result that earlier “Flash” releases did not produce.
Why it’s interesting
Section titled “Why it’s interesting”This is the Flash-tier shoe to drop after Nano Banana 2 / Gemini 3.1 Flash Image (Google product announcement) (Nano Banana 2 / Gemini 3.1 Flash Image) and Gemini 3.1 Flash Live: realtime voice and vision agent model (Google product announcement) (Gemini 3.1 Flash Live) — the same model family, now with the reasoning + agentic text checkpoint shipped openly. Together they fill in what Sundar Pichai 🤔🤔 reply teasing Gemini 3 release timing teased and what Gemini 3 Deep Think: Advancing science, research and engineering announced on the Pro side.
The benchmark line that matters for Luma research is the 78% SWE-Bench Verified at Flash-tier pricing/latency: this is the same agentic-coding axis where the Open foundation-model releases concept page tracks Kimi K2.5 reporting frontier parity (60.6–78.4 BrowseComp) against closed Gemini 3 Pro. A Flash-tier Gemini now beating its own Pro tier on SWE-Bench complicates that comparison — it suggests the “frontier coding model” target is moving inside Google’s own lineup faster than the open-source benchmark cycle.
The thinking-level + thought-signatures plumbing is also worth noting: it’s the most explicit productization yet of Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities-style adjustable reasoning budgets, exposed as a developer knob alongside an Interactions API for multi-turn agent state.
See also
Section titled “See also”- Sundar Pichai 🤔🤔 reply teasing Gemini 3 release timing — the CEO-level teaser preceding the Gemini 3 launch wave; this is the Flash member of that family
- Gemini 3 Deep Think: Advancing science, research and engineering — Gemini 3 Deep Think product post on the Pro side of the same family
- Nano Banana 2 / Gemini 3.1 Flash Image (Google product announcement) — Nano Banana 2 / Gemini 3.1 Flash Image, the image counterpart in the same family
- Gemini 3.1 Flash Live: realtime voice and vision agent model (Google product announcement) — Gemini 3.1 Flash Live, the realtime voice+vision counterpart
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities — Gemini 2.5 technical report; predecessor generation with the original thinking-budget design
- Open foundation-model releases — Gemini 3 Flash is the closed-frontier counterpart to the open releases tracked there
- LLM Inference Efficiency — the cost/speed/quality Pareto-frontier framing this post leans on
- Agentic Software Engineering — 78% SWE-Bench Verified at Flash pricing is the data point for this concept