Skip to content

Qwen3-TTS (version 2025-11-27) announcement — 49 voices, 10 languages, 8 Chinese dialects, ~97ms first-packet latency

Qwen announces a new dated revision of Qwen3-TTS (version 2025-11-27), accessible via Qwen Chat’s read-aloud, an Alibaba Cloud Model Studio realtime API, an offline API, and HF/ModelScope demos. The headline numbers: over 49 voices, 10 languages (zh / en / de / it / pt / es / ja / ko / fr / ru) plus 8 authentic Chinese dialects (Minnan, Wu, Cantonese, Sichuan, Beijing, Nanjing, Tianjin, Shaanxi). No tech report or benchmark numbers are attached to this tweet — it is an API/demo refresh that precedes the January 2026 open-weights release of the same family (Qwen3-TTS Family is Now Open Sourced: Voice Design, Clone, and Generation).

  • Version is dated 2025-11-27, with qwen3-tts-flash-realtime-2025-11-27 and qwen3-tts-flash-2025-11-27 exposed as the realtime and offline Model Studio model IDs [tweet body].
  • “Over 49 high-quality voices” are claimed, spanning personality ranges (the tweet mentions “cute and playful to wise and stern” as anchors) [tweet body].
  • 10 languages are supported: zh, en, de, it, pt, es, ja, ko, fr, ru [tweet body].
  • 8 authentic Chinese dialects are supported: Minnan, Wu, Cantonese, Sichuan, Beijing, Nanjing, Tianjin, Shaanxi [tweet body].
  • Rhythm and speed are claimed to adapt “just like a real person” — no quantitative latency / WER numbers in this tweet [tweet body].
  • Access surfaces named: Qwen Chat read-aloud, Model Studio realtime API, Model Studio offline API, HF Spaces demo, ModelScope demo, and a qwen.ai/blog?id=qwen3-tts-1128 blog post [tweet body].

The tweet does not describe the architecture; it is a release announcement only. By backref to the open-weights variant (Qwen3-TTS Family is Now Open Sourced: Voice Design, Clone, and Generation) the family is a discrete-multi-codebook speech LM over a self-developed 12 Hz acoustic tokenizer, with task-specialized heads (Base for cloning, VoiceDesign for description-driven synthesis, CustomVoice for preset-timbre style control) and a “dual-track hybrid streaming” architecture quoted at ~97 ms first-packet latency in the later open release. Whether the 2025-11-27 closed version uses the same backbone, tokenizer, and head split is not stated in the tweet — the open release happens roughly seven weeks later and adds the “12Hz” tokenizer name + 0.6B/1.7B size split publicly.

No benchmark numbers in this announcement. The tweet leans on qualitative claims (“uncanny”, “insanely natural”) and capability counts (49+ voices, 10 languages, 8 dialects). Engagement at the time of fetch: 184.2K views, 65 replies, 279 retweets, 2K likes.

This is the dated API/closed-version that lands between the open Qwen3-TTS family the wiki already has filed and any future open dump of the 2025-11-27 weights — useful as a pinning datapoint for the release timeline. Compared to Qwen3-TTS Family is Now Open Sourced: Voice Design, Clone, and Generation (open-source, dated qwen3tts-0115, 6 checkpoints across 2 sizes, full tech report at arXiv:2601.15621), this tweet covers a predecessor closed revision with a wider claimed voice/dialect surface (49 voices + 8 Chinese dialects vs. the open release’s 9 preset CustomVoice timbres) but no architectural disclosure and no open weights. The 8-Chinese-dialect line in particular is more explicit here than in the open-release blog and is likely the marketing wedge for the Chinese consumer-product use cases. Sits next to Gemini 3.1 Flash TTS: Google text-to-speech with scene direction and audio tags (announcement) (Gemini 3.1 Flash TTS) as part of the late-2025 frontier-lab TTS push that the joint audio-video releases (LTX-2: Efficient Joint Audio-Visual Foundation Model, Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation) are partially competing with on the audio side.