Qwen3-TTS (version 2025-11-27) announcement — 49 voices, 10 languages, 8 Chinese dialects, ~97ms first-packet latency
Qwen announces a new dated revision of Qwen3-TTS (version 2025-11-27), accessible via Qwen Chat’s read-aloud, an Alibaba Cloud Model Studio realtime API, an offline API, and HF/ModelScope demos. The headline numbers: over 49 voices, 10 languages (zh / en / de / it / pt / es / ja / ko / fr / ru) plus 8 authentic Chinese dialects (Minnan, Wu, Cantonese, Sichuan, Beijing, Nanjing, Tianjin, Shaanxi). No tech report or benchmark numbers are attached to this tweet — it is an API/demo refresh that precedes the January 2026 open-weights release of the same family (Qwen3-TTS Family is Now Open Sourced: Voice Design, Clone, and Generation).
Key claims
Section titled “Key claims”- Version is dated
2025-11-27, withqwen3-tts-flash-realtime-2025-11-27andqwen3-tts-flash-2025-11-27exposed as the realtime and offline Model Studio model IDs [tweet body]. - “Over 49 high-quality voices” are claimed, spanning personality ranges (the tweet mentions “cute and playful to wise and stern” as anchors) [tweet body].
- 10 languages are supported: zh, en, de, it, pt, es, ja, ko, fr, ru [tweet body].
- 8 authentic Chinese dialects are supported: Minnan, Wu, Cantonese, Sichuan, Beijing, Nanjing, Tianjin, Shaanxi [tweet body].
- Rhythm and speed are claimed to adapt “just like a real person” — no quantitative latency / WER numbers in this tweet [tweet body].
- Access surfaces named: Qwen Chat read-aloud, Model Studio realtime API, Model Studio offline API, HF Spaces demo, ModelScope demo, and a
qwen.ai/blog?id=qwen3-tts-1128blog post [tweet body].
Method
Section titled “Method”The tweet does not describe the architecture; it is a release announcement only. By backref to the open-weights variant (Qwen3-TTS Family is Now Open Sourced: Voice Design, Clone, and Generation) the family is a discrete-multi-codebook speech LM over a self-developed 12 Hz acoustic tokenizer, with task-specialized heads (Base for cloning, VoiceDesign for description-driven synthesis, CustomVoice for preset-timbre style control) and a “dual-track hybrid streaming” architecture quoted at ~97 ms first-packet latency in the later open release. Whether the 2025-11-27 closed version uses the same backbone, tokenizer, and head split is not stated in the tweet — the open release happens roughly seven weeks later and adds the “12Hz” tokenizer name + 0.6B/1.7B size split publicly.
Results
Section titled “Results”No benchmark numbers in this announcement. The tweet leans on qualitative claims (“uncanny”, “insanely natural”) and capability counts (49+ voices, 10 languages, 8 dialects). Engagement at the time of fetch: 184.2K views, 65 replies, 279 retweets, 2K likes.
Why it’s interesting
Section titled “Why it’s interesting”This is the dated API/closed-version that lands between the open Qwen3-TTS family the wiki already has filed and any future open dump of the 2025-11-27 weights — useful as a pinning datapoint for the release timeline. Compared to Qwen3-TTS Family is Now Open Sourced: Voice Design, Clone, and Generation (open-source, dated qwen3tts-0115, 6 checkpoints across 2 sizes, full tech report at arXiv:2601.15621), this tweet covers a predecessor closed revision with a wider claimed voice/dialect surface (49 voices + 8 Chinese dialects vs. the open release’s 9 preset CustomVoice timbres) but no architectural disclosure and no open weights. The 8-Chinese-dialect line in particular is more explicit here than in the open-release blog and is likely the marketing wedge for the Chinese consumer-product use cases. Sits next to Gemini 3.1 Flash TTS: Google text-to-speech with scene direction and audio tags (announcement) (Gemini 3.1 Flash TTS) as part of the late-2025 frontier-lab TTS push that the joint audio-video releases (LTX-2: Efficient Joint Audio-Visual Foundation Model, Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation) are partially competing with on the audio side.
See also
Section titled “See also”- Qwen3-TTS Family is Now Open Sourced: Voice Design, Clone, and Generation — the open-weights successor of the same family (~7 weeks later), with full tech report, vLLM-Omni day-0 serving, and the Base / CustomVoice / VoiceDesign head split made public
- Open foundation-model releases — Qwen3-TTS as a single-model-series open-release pattern, of which this tweet is the closed API precursor
- Gemini 3.1 Flash TTS: Google text-to-speech with scene direction and audio tags (announcement) — Gemini 3.1 Flash TTS announcement, a contemporaneous closed-frontier TTS counterpart
- LTX-2: Efficient Joint Audio-Visual Foundation Model — sibling open audio model (joint AV), architectural foil (MMDiT + audio branch vs. discrete-LM + non-DiT decoder)
- Blog: https://qwen.ai/blog?id=qwen3-tts-1128 (link in tweet; body not retrievable at filing time)
- HF demo: https://huggingface.co/spaces/Qwen/Qwen3-TTS-Demo
- ModelScope demo: https://modelscope.cn/studios/Qwen/Qwen3-TTS-Demo
- Realtime API:
qwen3-tts-flash-realtime-2025-11-27(Alibaba Cloud Model Studio) - Offline API:
qwen3-tts-flash-2025-11-27(Alibaba Cloud Model Studio)