Skip to content

Wan 2.6 — native multimodal video model with audio, multi-shot storytelling, starring references (Alibaba Wan announcement)

Alibaba’s Wan team announces Wan 2.6, a native multimodal video model that generates synchronized audio with video in a single pass. Headline features: a “Starring” reference-to-video system that casts characters from reference videos into new scenes (R2V), native audio with phoneme-level lip sync (V2A), and Smart Multi-Shot, which decomposes a long narrative prompt into multiple shots with planned transitions and consistent characters. Reported envelope: up to 15-second 1080p clips, multi-aspect ratios, English and Chinese, on a 14B MoE backbone descended from Wan 2.2’s architecture. The release is positioned as a closed-source datapoint alongside Veo 3, Sora 2, and Kling 2.6 in the native joint-audio-video cluster; whether weights ship open like Wan 2.1/2.2 was not stated in the announcement tweet.

  • Wan 2.6 is positioned as a “native multimodal” model with three flagship features: Starring (reference-to-video character casting), native audio-video generation, and Smart Multi-Shot narrative decomposition [tweet body].
  • Starring supports human or human-like figures and is explicitly framed for “complex multi-person and human-object interactions” — i.e. it goes beyond single-subject reference conditioning [tweet body].
  • Native audio in Wan 2.6 is a single-pass joint generation, not a T2V→V2A post-process — the model produces dialogue with lip sync, sound effects, and ambient sound directly from frames [secondary coverage].
  • Smart Multi-Shot automatically segments a long prompt into multiple shots with planned transitions while maintaining character, lighting, and scene consistency across shots — a narrative-storytelling capability, not just longer single clips [secondary coverage].
  • Reported output envelope: up to 15-second clips at 1080p with multiple aspect ratios (16:9, 9:16, 4:3, 3:4, 1:1) [secondary coverage].

The tweet itself is a launch announcement with no architectural disclosure; technical details come from secondary product-page coverage of the API. The reported lineage is the open Wan 2.2 MoE backbone (14B activated), now extended with a native audio pathway. The Starring system uses reference video(s) as additional conditioning to preserve identity across generated shots. Multi-Shot is described as automatic prompt-to-shot-list decomposition with cross-shot consistency on characters, environment, and lighting. The model accepts text-to-video, image-to-video, and reference-to-video inputs. As with the Kling 2.6 announcement (Kling Video 2.6 — first Kling AI model with native audio (announcement)), no technical report, parameter counts beyond MoE 14B, audio sample rate, or benchmark numbers are disclosed in the launch material itself.

No benchmark numbers in the tweet. Secondary coverage cites comparisons against Veo 3.1, Sora 2, and Kling 2.6/3.0 as positioning peers, and Wan’s earlier 2.5 (with audio support already in preview at 10 s × 1080p × 24 FPS) as the immediate predecessor. The headline product claim is 15-second multi-shot 1080p with native synchronized audio in one pass — a duration and structural step beyond Wan 2.5’s 10 s single-shot envelope.