Wan 2.6 — native multimodal video model with audio, multi-shot storytelling, starring references (Alibaba Wan announcement)
Alibaba’s Wan team announces Wan 2.6, a native multimodal video model that generates synchronized audio with video in a single pass. Headline features: a “Starring” reference-to-video system that casts characters from reference videos into new scenes (R2V), native audio with phoneme-level lip sync (V2A), and Smart Multi-Shot, which decomposes a long narrative prompt into multiple shots with planned transitions and consistent characters. Reported envelope: up to 15-second 1080p clips, multi-aspect ratios, English and Chinese, on a 14B MoE backbone descended from Wan 2.2’s architecture. The release is positioned as a closed-source datapoint alongside Veo 3, Sora 2, and Kling 2.6 in the native joint-audio-video cluster; whether weights ship open like Wan 2.1/2.2 was not stated in the announcement tweet.
Key claims
Section titled “Key claims”- Wan 2.6 is positioned as a “native multimodal” model with three flagship features: Starring (reference-to-video character casting), native audio-video generation, and Smart Multi-Shot narrative decomposition [tweet body].
- Starring supports human or human-like figures and is explicitly framed for “complex multi-person and human-object interactions” — i.e. it goes beyond single-subject reference conditioning [tweet body].
- Native audio in Wan 2.6 is a single-pass joint generation, not a T2V→V2A post-process — the model produces dialogue with lip sync, sound effects, and ambient sound directly from frames [secondary coverage].
- Smart Multi-Shot automatically segments a long prompt into multiple shots with planned transitions while maintaining character, lighting, and scene consistency across shots — a narrative-storytelling capability, not just longer single clips [secondary coverage].
- Reported output envelope: up to 15-second clips at 1080p with multiple aspect ratios (16:9, 9:16, 4:3, 3:4, 1:1) [secondary coverage].
Method
Section titled “Method”The tweet itself is a launch announcement with no architectural disclosure; technical details come from secondary product-page coverage of the API. The reported lineage is the open Wan 2.2 MoE backbone (14B activated), now extended with a native audio pathway. The Starring system uses reference video(s) as additional conditioning to preserve identity across generated shots. Multi-Shot is described as automatic prompt-to-shot-list decomposition with cross-shot consistency on characters, environment, and lighting. The model accepts text-to-video, image-to-video, and reference-to-video inputs. As with the Kling 2.6 announcement (Kling Video 2.6 — first Kling AI model with native audio (announcement)), no technical report, parameter counts beyond MoE 14B, audio sample rate, or benchmark numbers are disclosed in the launch material itself.
Results
Section titled “Results”No benchmark numbers in the tweet. Secondary coverage cites comparisons against Veo 3.1, Sora 2, and Kling 2.6/3.0 as positioning peers, and Wan’s earlier 2.5 (with audio support already in preview at 10 s × 1080p × 24 FPS) as the immediate predecessor. The headline product claim is 15-second multi-shot 1080p with native synchronized audio in one pass — a duration and structural step beyond Wan 2.5’s 10 s single-shot envelope.
Why it’s interesting
Section titled “Why it’s interesting”- This is the third major closed/proprietary launch in the native joint-audio-video cluster within a few weeks, after Kling 2.6 (Kling Video 2.6 — first Kling AI model with native audio (announcement)) and the Veo 3 / Sora 2 lineage (Veo 3 Tech Report). The convergence is striking — every frontier video lab is shipping native audio in late 2025 — but technical disclosure remains sparse on the closed side, in contrast to the open-source dual-stream-DiT recipe converged on by LTX-2 (LTX-2: Efficient Joint Audio-Visual Foundation Model), MOVA (MOVA: Towards Scalable and Synchronized Video-Audio Generation), and SkyReels-V4 (SkyReels-V4: Multi-modal Video-Audio Generation, Inpainting and Editing model).
- Multi-shot narrative decomposition with cross-shot character/lighting consistency is a meaningfully different framing from the dominant academic recipe (single-shot up-to-N-seconds joint A+V), and overlaps with the territory of ShotAdapter (ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models) and TalkCuts (TalkCuts: A Large-Scale Dataset for Multi-Shot Human Speech Video Generation) — Alibaba is integrating it into the base model rather than as an adapter or downstream dataset.
- The Starring R2V system continues the line of identity-preserving multi-reference video generation already explored by Wan-Animate (Wan-Animate: Unified Character Animation and Replacement with Holistic Replication) and Phantom-Data (Phantom-Data: Towards a General Subject-Consistent Video Generation Dataset), now bundled into the same base model as the audio pathway. Worth watching whether weights drop open like prior Wan releases (Wan 2.1, 2.2 were Apache 2.0), or whether Wan 2.6 is the line where Alibaba retains the recipe.
See also
Section titled “See also”- Joint audio-video generation — third closed-source datapoint after Kling 2.6, Veo 3, Sora 2
- Kling Video 2.6 — first Kling AI model with native audio (announcement) — direct closed-source peer (Kling 2.6 native audio launch)
- LTX-2: Efficient Joint Audio-Visual Foundation Model — open dual-stream T2AV baseline (LTX-2)
- MOVA: Towards Scalable and Synchronized Video-Audio Generation — open MoE T2AV peer (MOVA)
- SkyReels-V4: Multi-modal Video-Audio Generation, Inpainting and Editing model — open multi-shot T2AV peer (SkyReels-V4, also targets 15 s × 1080p × multi-shot)
- Wan2.2: World's First Open-Source MoE Video Generation Model (launch announcement) — Wan 2.2 launch (the open MoE predecessor backbone)
- Wan 2.5 live on WaveSpeed — 5s/10s clips up to 1080p (AI Pulse correction tweet) — Wan 2.5 availability on WaveSpeed (immediate predecessor with preview audio)
- Wan-Animate: Unified Character Animation and Replacement with Holistic Replication — Wan-Animate (the identity-preservation lineage Starring extends)
- TalkCuts: A Large-Scale Dataset for Multi-Shot Human Speech Video Generation — multi-shot human-speech video dataset (overlapping territory)
- ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models — adapter approach to multi-shot text-to-video (contrasting approach)