Skip to content

Qwen3.8 open weights release — 27B dense multimodal + 2.4T-A95B MoE (Alibaba Qwen)

Alibaba Qwen’s August 14, 2026 announcement of the Qwen3.8 open-weights release — the follow-through on the July 19 preview tease (Qwen3.8 preview announcement — 2.4T-parameter frontier model, open-weights soon (Alibaba Qwen)). Two tiers ship together: (1) Qwen3.8-27B, a native multimodal dense model with 262K native context extendable to 1M via YaRN, positioned as beating Qwen3.7-Plus overall while excelling in coding + office workflows, and (2) Qwen3.8-2.4T-A95B, the Max-tier MoE (2.4T total / 95B active) whose open weights the tweet notes “have also been released recently.” Both are Apache 2.0 on Hugging Face and ModelScope; day-0 Unsloth Dynamic GGUFs land alongside for local inference and fine-tuning.

  • Qwen3.8-27B is a 27B dense native-multimodal model licensed Apache 2.0, claimed to beat Qwen3.7-Plus on overall benchmarks and to shine on real-world coding + office workflows [tweet body].
  • Context length: 262K native, extendable to 1M tokens via YaRN [tweet body].
  • Qwen3.8-2.4T-A95B (Max tier) open weights have also been released, per the launch tweet — 2.4T total parameters with ~95B active in the MoE, matching the July 19 preview’s total-parameter count [tweet body; cf. Qwen3.8 preview announcement — 2.4T-parameter frontier model, open-weights soon (Alibaba Qwen)].
  • Performance chart embedded in the follow-up tweet (image not fetched) is the sole benchmark disclosure; no per-benchmark numbers are transcribed in the announcement text itself [tweet reply 1].
  • Distribution: Hugging Face collection at huggingface.co/collections/Qwen/qwen38 and ModelScope collection at modelscope.cn/collections/Qwen/Qwen38 [tweet body embedded URLs].
  • Unsloth ships Dynamic GGUFs for Qwen3.8-27B at unsloth/Qwen3.8-27B-GGUF with day-0 Unsloth Desktop support for running and fine-tuning [reply from @UnslothAI].

Product-launch tweet — no architecture paper, no training-recipe details, no benchmark table transcribed. Only quantitative claims are the tier configurations (27B dense multimodal; 2.4T total / 95B active MoE) and the context lengths (262K native, 1M via YaRN). The 27B tier is notable as dense rather than MoE — Qwen has been shipping sparse-MoE flagships since Qwen3-Next-80B-A3B, so a dense 27B in the Qwen3.8 lineup echoes the Qwen3.6-27B dense-beats-larger-MoE result the wiki already tracks (Qwen3.6-27B announcement: 27B dense beats Qwen3.5-397B-A17B on coding benchmarks) — a 27B dense that outperforms Qwen3.5-397B-A17B on coding benchmarks.

The staggered release shape matches the Qwen pattern: closed API preview first (Jul 19), then open weights across dense + MoE tiers (Aug 14). Notable that both tiers ship together in one announcement rather than the two-step “small open first, flagship later” pattern seen elsewhere in the open-frontier cohort — the Max-tier 2.4T-A95B is disclosed as already open at the same moment as the 27B tier.

No benchmark numbers in the tweet text. Engagement at filing: 58.3K views on the launch tweet plus 15K views on the “Performance of Qwen3.8-27B:” reply with two benchmark screenshots (image content not fetched here — needs a follow-up read on the HF collection or ModelScope card for actual numbers). The Unsloth reply adds the day-0 GGUF quantizations + Unsloth Desktop integration.

This is the concrete open-weights follow-through on the Qwen3.8 preview announcement — 2.4T-parameter frontier model, open-weights soon (Alibaba Qwen) tease, and it lands in the same week as Alibaba’s other open-cadence hits (Gemini 3.7 Flash — 50% cheaper than 3.6, ~3 weeks between releases on Gemini 3.7 Flash’s closed algorithmic gains highlights the contrast). Two things make it notable beyond routine Qwen versioning:

Complements Qwen3-TTS Family is Now Open Sourced: Voice Design, Clone, and Generation and the Qwen3-VL-Embedding+Reranker joint release (Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking) as another Qwen “series shipped as one coordinated release” datapoint.