Gemini Native Audio — updated model with higher-precision function calling, better realtime instruction following, smoother conversation
Logan Kilpatrick (Google AI Studio) announces an updated Gemini Native Audio model, available in the Gemini API. The pitch lists three improvements: higher-precision function calling, better realtime instruction following, and smoother and more cohesive conversation. No technical report, no benchmark numbers, no architecture details, no model card link — this is a closed, API-only product announcement. Filed as a tracking pointer for Google’s realtime/streaming audio stack and as another Gemini-3.1-Flash-era closed-frontier datapoint.
Key claims
Section titled “Key claims”- The updated Gemini Native Audio model has higher-precision function calling than the prior version [tweet body].
- The model has better realtime instruction following than the prior version [tweet body].
- Conversational abilities are described as smoother and more cohesive than the prior version [tweet body].
- Delivery is via the Gemini API, available to developers immediately [tweet body].
Method
Section titled “Method”Not disclosed. The tweet contains no architecture, no parameter count, no audio tokenizer or codec detail, no training data description, no latency or quality numbers, and no model identifier in the post body itself. External signals at filing time indicate the underlying serving model is gemini-2.5-flash-native-audio-preview-12-2025 (this is inference from Google’s GitHub bug-tracker references, not stated in the tweet). Naming places “Native Audio” alongside the realtime sibling Gemini 3.1 Flash Live: realtime voice and vision agent model (Google product announcement) (Gemini 3.1 Flash Live) and the TTS sibling Gemini 3.1 Flash TTS: Google text-to-speech with scene direction and audio tags (announcement) (Gemini 3.1 Flash TTS), with Native Audio appearing to be the in-API voice-output endpoint that powers the Live realtime mode rather than a separate model lineage — but this is inference, not claimed.
Results
Section titled “Results”None reported. No measured first-packet latency, no function-call accuracy numbers, no instruction-following benchmark, no MOS or naturalness rating, no head-to-head against OpenAI Realtime API or any open-weight realtime voice model. The only quantitative signal on the page is the 115K view count, which is irrelevant to capability.
Why it’s interesting
Section titled “Why it’s interesting”This is the earlier Native Audio rollout in Google’s late-2025 / early-2026 closed multimodal speech cadence — it predates both Gemini 3.1 Flash Live: realtime voice and vision agent model (Google product announcement) (Gemini 3.1 Flash Live, realtime voice + vision agent) and Gemini 3.1 Flash TTS: Google text-to-speech with scene direction and audio tags (announcement) (Gemini 3.1 Flash TTS, scene-direction + audio tags + 70 languages). The three explicit improvements named — function calling precision, realtime instruction following, conversational cohesion — are the dimensions where Gemini’s prior native-audio dialog endpoint was reportedly weakest in the developer-issue threads, so this is the model update positioning itself against those known failure modes. Combined with the later siblings, it completes a four-quarter picture of Google shipping the Gemini-3.1 voice stack one capability at a time via @OfficialLoganK product tweets: Native Audio dialog → Flash Live (vision-grounded realtime) → Flash TTS (controllable expressive synthesis). For the Open foundation-model releases cohort, it adds another closed-but-API-accessible baseline that open realtime voice models — most directly PersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models (PersonaPlex full-duplex), Voxtral TTS — Mistral AI frontier open-weight text-to-speech model (announcement) (Voxtral TTS), and Qwen3-TTS Family is Now Open Sourced: Voice Design, Clone, and Generation (Qwen3-TTS family with concrete WER + 97ms first-packet latency) — get measured against.
See also
Section titled “See also”- Gemini 3.1 Flash Live: realtime voice and vision agent model (Google product announcement) — same author, same product line, realtime voice + vision sibling (Gemini 3.1 Flash Live) that builds on top of Native Audio
- Gemini 3.1 Flash TTS: Google text-to-speech with scene direction and audio tags (announcement) — same author, same product line, controllable TTS sibling (Gemini 3.1 Flash TTS)
- Nano Banana 2 / Gemini 3.1 Flash Image (Google product announcement) — same author, same Gemini 3.1 Flash family, image-gen sibling (Nano Banana 2)
- Gemini Embedding 2: SOTA multimodal embedding model (Google product announcement) — same author, multimodal embedding sibling in the same rollout
- PersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models — closest concurrent realtime full-duplex conversational-speech research datapoint
- Qwen3-TTS Family is Now Open Sourced: Voice Design, Clone, and Generation — open-weight TTS counterpart with concrete WER + latency numbers
- Voxtral TTS — Mistral AI frontier open-weight text-to-speech model (announcement) — direct open-weight TTS counterpart (Voxtral TTS)
- CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models — open streaming TTS technical reference for the kind of architecture this kind of closed product likely sits on
- Open foundation-model releases — Gemini Native Audio adds another datapoint to the closed-but-API-accessible cohort that open releases position against