Skip to content

Voxtral TTS — Mistral AI frontier open-weight text-to-speech model (announcement)

Mistral AI’s tweet announces Voxtral TTS, a new open-weight text-to-speech model framed as their frontier TTS offering. Headline claims: emotionally expressive natural speech, 9-language coverage with diverse dialects, very low time-to-first-audio latency, and easy adaptation to new voices. Mistral reports that in zero-shot custom-voice tests judged by native speakers (naturalness, accent accuracy, similarity to original), Voxtral TTS beats ElevenLabs v2.5 Flash. The model is available via the Mistral Studio playground and pairs with Voxtral Transcribe for end-to-end speech-to-speech pipelines.

  • Voxtral TTS is positioned as a new frontier open-weight TTS model emphasizing natural and expressive speech [tweet OP].
  • Supports 9 languages with accurate dialect rendering, targeting global voice workflows [tweet OP].
  • Low time-to-first-audio latency, advertised as suitable for real-time use cases like customer support and live translation [tweet OP].
  • Designed to compose with Voxtral Transcribe for end-to-end speech-to-speech, or to drop into an existing STT + LLM stack [tweet OP].
  • In zero-shot custom-voice tests judged by native speakers, Voxtral TTS reportedly outperforms ElevenLabs v2.5 Flash on naturalness, accent accuracy, and similarity to the original voice [tweet OP].
  • Available immediately in the Mistral Studio playground with both Mistral preset voices and user-recorded reference voices [tweet OP].

Architectural details are not disclosed in the tweet. The announcement positions Voxtral TTS as the output-side companion to the previously released Voxtral (Transcribe) ASR family, suggesting the broader “Voxtral” naming covers both directions of the speech stack. The tweet thread also references adjacent Mistral releases (Devstral 2 coding models, Mistral Vibe CLI, the original Voxtral ASR models built on the Mistral Small 3.1 backbone with 32k context), but Voxtral TTS itself is described purely by its capabilities — voice cloning, zero-shot custom voice from user recordings, instruction-style adaptability — with no architecture, parameter count, tokenizer, or training-data details.

The only quantitative claim is the head-to-head comparison: in zero-shot custom-voice tests, native-speaker judges rated Voxtral TTS above ElevenLabs v2.5 Flash on naturalness, accent accuracy, and similarity to the original voice. No WER numbers, no MOS scores, no latency figures with hardware, and no benchmark table are included in the tweet itself.

Voxtral TTS is the direct open-weight counterpart to ElevenLabs in the same way Qwen3-TTS Family is Now Open Sourced: Voice Design, Clone, and Generation is — both announce zero-shot voice cloning, multi-language coverage, and head-to-head wins against closed TTS APIs as the headline pitch. Tracking which open TTS family ends up the strongest Luma-usable baseline for character voices and dubbing matters; Qwen3-TTS reports 0.77/1.24 WER on Seed-TTS at 1.7B with 97 ms first-packet latency, and Voxtral now claims a similar “beats closed-API” framing without releasing the numbers yet. Worth waiting for a tech report or model card before treating the ElevenLabs-beat claim as more than marketing.

It also fits the Open foundation-model releases pattern that Mistral has been running on the LLM side — see Mistral Small 4 119B (instruct + reasoning + Devstral unified MoE) for the unified Mistral Small 4 release (instruct + reasoning + Devstral in one MoE). Same playbook: announce a frontier-level open-weight artifact paired with playground access and a coordinated companion (Voxtral Transcribe for speech-to-speech, Mistral Vibe CLI for the coding family).