Skip to content

Gemini 3.1 Flash Live: realtime voice and vision agent model (Google product announcement)

Logan Kilpatrick (Google AI Studio) announces Gemini 3.1 Flash Live, Google’s new realtime model for building voice and vision agents. The post claims “more than a year” of work across model, infrastructure, and product experience, and frames the release as a “step function improvement in quality, reliability, and latency” over prior realtime offerings. No technical report, no benchmark numbers, no architecture details — this is a closed, API-only product announcement. Filed as a tracking pointer for Google’s realtime/streaming multimodal stack and as a closed-frontier comparison baseline for any future open realtime voice+vision work.

  • Gemini 3.1 Flash Live is positioned as a realtime model intended for building voice and vision agents [tweet body].
  • Google claims the release reflects more than a year of joint work on model, infra, and experience [tweet body].
  • The pitch is a “step function improvement” along three axes simultaneously: quality, reliability, and latency [tweet body].

Not disclosed. The tweet contains no architecture, training data, parameter count, modality-fusion design, audio tokenizer detail, or latency numbers. Naming places it in the Gemini 3.1 Flash family — the same family as the closed image-generation model in Nano Banana 2 / Gemini 3.1 Flash Image (Google product announcement) (Nano Banana 2 / Gemini 3.1 Flash Image) — suggesting a shared Flash-class backbone routed to a realtime streaming head rather than a separate model lineage. Two attached product screenshots are referenced but not parsed here.

None reported. No measured latency, no audio quality numbers, no vision benchmarks, no comparison table against prior realtime offerings (OpenAI Realtime API, Gemini Live earlier versions). The only quantitative signal on the page is the 315.5K view count, which is irrelevant to capability.

This is the realtime/streaming counterpart to the still-image and text Gemini 3.x Flash releases the wiki already tracks — Nano Banana 2 / Gemini 3.1 Flash Image (Google product announcement) (Nano Banana 2 / Gemini 3.1 Flash Image), Gemini 3 Deep Think: Advancing science, research and engineering (reasoning), and Introducing Agentic Vision in Gemini 3 Flash (agentic vision). The “voice + vision agent” framing is the closed-API datapoint that Luma’s interactive-video work — Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation, Xmax X1 — Real-Time Interactive Video Model (Product Announcement), Helios: Real Real-Time Long Video Generation Model — will get benchmarked against on the perception side once any concrete latency or capability numbers surface. Filing the announcement gives those future open releases a stable pointer for what “Gemini Live” actually was at the time of comparison.