Skip to content

Qwen3-Omni-Flash 2025-12-01 update — multi-turn audio/video, system-prompt persona, 119 text / 19 speech languages

Alibaba Qwen ships a 2025-12-01 update to Qwen3-Omni-Flash, the predecessor-generation omni-modal model (text + image + audio + video in, text + speech out). The update headlines four changes: enhanced multi-turn video/audio understanding, system-prompt-controllable AI personality (roleplay), expanded language coverage at 119 text languages / 19 speech languages, and TTS voices the announcement claims are indistinguishable from humans. Distributed through Qwen Chat’s VoiceChat/VideoChat buttons, HF + ModelScope demo spaces, and Alibaba Cloud realtime/offline APIs.

  • Qwen3-Omni-Flash 2025-12-01 enhances multi-turn video and audio understanding so conversations “flow naturally” [tweet body].
  • System-prompt-driven persona customization (roleplay) is exposed for the Flash variant in this update [tweet body].
  • Language coverage at this release: 119 text languages and 19 speech languages [tweet body].
  • TTS voice quality is claimed to be “indistinguishable from humans” — no benchmark or MOS reported in the tweet [tweet body].
  • Distribution channels: Qwen Chat (VoiceChat / VideoChat), HF Space Qwen/Qwen3-Omni-Flash, ModelScope Studio, plus Alibaba Cloud Model Studio realtime and offline APIs [tweet body, linked endpoints].

Announcement-only; no architecture details in the tweet. The linked blog (qwen.ai/blog?id=qwen3-omni-flash) was not retrievable at filing time. Based on the predecessor lineage, Qwen3-Omni-Flash is the smaller-tier sibling of the Qwen3-Omni Plus omni-modal model (Thinker-Talker decoder split, audio encoder feeding a hybrid-attention backbone, streaming RVQ codec + ConvNet code2wav decoder for speech output) — the 2026-04 successor Qwen3.5-Omni technical report (Qwen3.5-Omni Technical Report) documents the architectural family in detail and reports Flash latency at 235 ms first-packet (audio) / 426 ms (video) at 1× concurrency. The 2025-12-01 update is presented as a model refresh on the existing serving stack, not a new architecture.

No benchmarks in the tweet. The headline numbers are surface counts:

  • 119 text languages, 19 speech languages [tweet body].
  • 127.3K tweet views at filing time (engagement signal, not a model metric).

This update sits exactly between two filed Qwen omni-modal releases: it post-dates the Qwen3-Omni launch family (covered by Qwen3-TTS Family is Now Open Sourced: Voice Design, Clone, and Generation‘s TTS branch) and pre-dates the architectural overhaul in Qwen3.5-Omni Technical Report / Qwen3.5-Omni: A Native Omni-Modal Model with Thinker-Talker MoE Architecture. It’s worth filing because the language-count expansion (19 speech languages) and the system-prompt persona handle indicate which features were judged stable enough to back-port to the Flash tier before the 3.5 jump, which is useful for tracking what Qwen treats as production-ready vs. experimental. It also extends the Open foundation-model releases cluster’s record of Qwen’s serial-update cadence — Qwen ships multiple in-place refreshes per generation rather than only major-version bumps.