Gemini 3.1 Flash TTS: Google text-to-speech with scene direction and audio tags (announcement)
Product-launch tweet from Logan Kilpatrick (Google AI Studio) announcing Gemini 3.1 Flash TTS, Google’s latest text-to-speech model. Headline capability claims: scene direction, speaker-level specificity, audio tags, more natural and expressive voices, and 70-language coverage. Closed and API-only — available through a new audio playground in AI Studio and via the Gemini API, with no technical report, no benchmark numbers, no weights. Filed as a tracking pointer for Google’s closed multimodal speech stack and as a closed-frontier comparison baseline for open TTS releases.
Key claims
Section titled “Key claims”- Gemini 3.1 Flash TTS is Google’s latest text-to-speech model, shipped under the same Gemini 3.1 Flash family naming used for image, embedding, and realtime variants [tweet body].
- The model supports scene direction and speaker-level specificity, plus audio tags as a control interface for expressive synthesis [tweet body].
- Voices are positioned as more natural and expressive than prior Google TTS offerings [tweet body].
- 70 languages are supported at launch [tweet body].
- Delivery is via a new audio playground in AI Studio and the Gemini API — no weights, no model card, no benchmarks in the announcement [tweet body].
Method
Section titled “Method”Not disclosed. The tweet contains no architecture, no parameter count, no audio tokenizer or codec detail, no training data description, no latency or quality numbers. Naming places it in the Gemini 3.1 Flash family alongside Nano Banana 2 / Gemini 3.1 Flash Image (Nano Banana 2 / Gemini 3.1 Flash Image (Google product announcement)) and Gemini 3.1 Flash Live (Gemini 3.1 Flash Live: realtime voice and vision agent model (Google product announcement)), which suggests a shared Flash-class backbone routed to a speech-synthesis head rather than a separate model lineage — but this is inference, not claimed. “Scene direction” and “audio tags” imply an instruction-style control interface over prosody, emotion, and acoustic environment; “speaker-level specificity” implies multi-speaker support, possibly with reference-voice conditioning, but the tweet does not specify. The attached video preview is referenced but not parsed here.
Results
Section titled “Results”None reported. No MOS, no naturalness rating, no WER on a back-transcription benchmark, no head-to-head against ElevenLabs / OpenAI TTS / Qwen3-TTS / Voxtral TTS, no latency figure with hardware. The only quantitative signal on the page is the 797K view count, which is irrelevant to capability.
Why it’s interesting
Section titled “Why it’s interesting”Gemini 3.1 Flash TTS is the closed-source TTS-frontier baseline the wiki’s open-weight TTS pages position against — most directly Voxtral TTS — Mistral AI frontier open-weight text-to-speech model (announcement) (Voxtral TTS, frontier open-weight TTS announced as beating ElevenLabs v2.5 Flash) and Qwen3-TTS Family is Now Open Sourced: Voice Design, Clone, and Generation (Qwen3-TTS family with concrete Seed-TTS WER numbers and 97 ms first-packet latency). It also completes the Gemini-3.1-Flash modality grid that Open foundation-model releases tracks as the closed-but-API-accessible cohort: Flash-Image (Nano Banana 2 / Gemini 3.1 Flash Image (Google product announcement)), Flash-Live realtime voice+vision (Gemini 3.1 Flash Live: realtime voice and vision agent model (Google product announcement)), Flash-Embedding (Gemini Embedding 2: SOTA multimodal embedding model (Google product announcement)), and now Flash-TTS — Google is shipping the Gemini-3.1 family one capability at a time via @OfficialLoganK product tweets, and the speech-output piece had been the visible gap.
The “scene direction” framing is notable: it pitches TTS as a controllable expressive-synthesis interface rather than a text-in / waveform-out function, which is the dimension Mistral’s Voxtral and Alibaba’s Qwen3-TTS-VoiceDesign also emphasize. Once any numbers surface, this is the announcement to point future open releases at as the comparison baseline.
See also
Section titled “See also”- Nano Banana 2 / Gemini 3.1 Flash Image (Google product announcement) — same Gemini 3.1 Flash family, image-gen sibling (Nano Banana 2)
- Gemini Embedding 2: SOTA multimodal embedding model (Google product announcement) — same author, multimodal embedding sibling in the same rollout
- Gemini 3.1 Flash Live: realtime voice and vision agent model (Google product announcement) — same author, realtime voice+vision sibling (Gemini 3.1 Flash Live)
- Voxtral TTS — Mistral AI frontier open-weight text-to-speech model (announcement) — direct open-weight TTS counterpart (Voxtral TTS, beats ElevenLabs v2.5 Flash by self-report)
- Qwen3-TTS Family is Now Open Sourced: Voice Design, Clone, and Generation — open-weight TTS counterpart with concrete WER + latency numbers
- Open foundation-model releases — Gemini 3.1 Flash TTS adds a TTS datapoint to the closed-but-API-accessible cohort
- PersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models — adjacent conversational-speech research datapoint (full-duplex, voice/role control)
- Microsoft AI ships MAI-Transcribe-1, MAI-Voice-1, and MAI-Image-2 — Mustafa Suleyman announcement — concurrent Microsoft TTS announcement (MAI-Voice-1)