Skip to content

Gemini Native Audio — updated model with higher-precision function calling, better realtime instruction following, smoother conversation

Logan Kilpatrick (Google AI Studio) announces an updated Gemini Native Audio model, available in the Gemini API. The pitch lists three improvements: higher-precision function calling, better realtime instruction following, and smoother and more cohesive conversation. No technical report, no benchmark numbers, no architecture details, no model card link — this is a closed, API-only product announcement. Filed as a tracking pointer for Google’s realtime/streaming audio stack and as another Gemini-3.1-Flash-era closed-frontier datapoint.

  • The updated Gemini Native Audio model has higher-precision function calling than the prior version [tweet body].
  • The model has better realtime instruction following than the prior version [tweet body].
  • Conversational abilities are described as smoother and more cohesive than the prior version [tweet body].
  • Delivery is via the Gemini API, available to developers immediately [tweet body].

Not disclosed. The tweet contains no architecture, no parameter count, no audio tokenizer or codec detail, no training data description, no latency or quality numbers, and no model identifier in the post body itself. External signals at filing time indicate the underlying serving model is gemini-2.5-flash-native-audio-preview-12-2025 (this is inference from Google’s GitHub bug-tracker references, not stated in the tweet). Naming places “Native Audio” alongside the realtime sibling Gemini 3.1 Flash Live: realtime voice and vision agent model (Google product announcement) (Gemini 3.1 Flash Live) and the TTS sibling Gemini 3.1 Flash TTS: Google text-to-speech with scene direction and audio tags (announcement) (Gemini 3.1 Flash TTS), with Native Audio appearing to be the in-API voice-output endpoint that powers the Live realtime mode rather than a separate model lineage — but this is inference, not claimed.

None reported. No measured first-packet latency, no function-call accuracy numbers, no instruction-following benchmark, no MOS or naturalness rating, no head-to-head against OpenAI Realtime API or any open-weight realtime voice model. The only quantitative signal on the page is the 115K view count, which is irrelevant to capability.

This is the earlier Native Audio rollout in Google’s late-2025 / early-2026 closed multimodal speech cadence — it predates both Gemini 3.1 Flash Live: realtime voice and vision agent model (Google product announcement) (Gemini 3.1 Flash Live, realtime voice + vision agent) and Gemini 3.1 Flash TTS: Google text-to-speech with scene direction and audio tags (announcement) (Gemini 3.1 Flash TTS, scene-direction + audio tags + 70 languages). The three explicit improvements named — function calling precision, realtime instruction following, conversational cohesion — are the dimensions where Gemini’s prior native-audio dialog endpoint was reportedly weakest in the developer-issue threads, so this is the model update positioning itself against those known failure modes. Combined with the later siblings, it completes a four-quarter picture of Google shipping the Gemini-3.1 voice stack one capability at a time via @OfficialLoganK product tweets: Native Audio dialog → Flash Live (vision-grounded realtime) → Flash TTS (controllable expressive synthesis). For the Open foundation-model releases cohort, it adds another closed-but-API-accessible baseline that open realtime voice models — most directly PersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models (PersonaPlex full-duplex), Voxtral TTS — Mistral AI frontier open-weight text-to-speech model (announcement) (Voxtral TTS), and Qwen3-TTS Family is Now Open Sourced: Voice Design, Clone, and Generation (Qwen3-TTS family with concrete WER + 97ms first-packet latency) — get measured against.