Seedance 2.0: Advancing Video Generation for World Complexity
Seedance 2.0 is ByteDance’s second-generation video-creation foundation model, packaged as a unified multimodal audio–video joint generation system that accepts text, image, audio, and video inputs and supports image-to-video, video extension, video editing, and multi-reference conditioning under one model. The arXiv submission is the team technical report attached to the Feb 2026 Doubao / Dreamina / CapCut launch; the marketing-level disclosure pegs it at #1 on the Artificial Analysis text-to-video Arena (Elo ~1269 at release), with 15-second multi-shot 1080p output and native synchronized dual-channel audio. The arXiv pull only exposes the title + ~200-author roster of the Seedance team at fetch time — the body, architecture, and ablation details are not present in the wiki’s fetched copy — so claims below are limited to what’s structurally inferable plus what ByteDance Seed and downstream coverage have stated publicly. It is the closed-source counterpart to LTX-2 / MOVA / SkyReels-V4 and the canonical Veo 3 / Sora 2 / Kling 3.0 / Vidu Q3 cohort referenced as an open question on Joint audio-video generation.
Key claims
Section titled “Key claims”- Seedance 2.0 is positioned as a unified multimodal audio–video joint generation model accepting four input modalities (text, image, audio, video) under one architecture rather than as separate T2V + V2A pipelines [Title + author affiliation block; underlying source: ByteDance Seed launch post linked from the abstract].
- The model is the technical report behind the Feb 2026 production launch on Doubao (China) and Dreamina / CapCut, listing the full Seedance team across ByteDance Seed and adjacent groups [Author list].
- Per the launch post linked from the abstract, headline capabilities include 15-second multi-shot output, dual-channel audio, image-conditioned generation, video extension, and video editing — i.e. the joint-A+V interface subsumes the editing surface, rather than treating it as a separate model [linked Seed launch post].
- Note on fetch coverage: only the arXiv title + author block were returned by the fetcher; the body of the technical report is not present in the wiki’s snapshot. The architectural specifics (capacity allocation, RoPE alignment between video and audio streams, multi-reference conditioning mechanism, distillation recipe) are referenced as “innovations in DiT temporal attention and variable-length generation” by secondary coverage but cannot be directly cited from the fetched text [arXiv page metadata only].
Method
Section titled “Method”Mechanically, the fetched arXiv copy does not expose method content beyond the title and the author roll. From the linked Seed launch and secondary technical coverage, Seedance 2.0 is described as a dual-branch (i.e. dual-stream) DiT that emits synchronized video and dual-channel audio in a single pass and accepts “comprehensive multi-reference inputs” (character, action template, cinematography style references plus audio). The model supports a multi-shot output mode and a video-editing/extension mode under the same backbone. This puts it structurally in the same family as the open dual-stream T2AV cluster (LTX-2: Efficient Joint Audio-Visual Foundation Model, MOVA: Towards Scalable and Synchronized Video-Audio Generation, SkyReels-V4: Multi-modal Video-Audio Generation, Inpainting and Editing model) but with closed weights and a closed training recipe.
Because the wiki’s fetched arXiv body is empty of architectural detail, the paper page deliberately stops at the interface-level description and does not enumerate technical contributions; a future /bud refresh once the full report or an HTML mirror is available is the right way to extend Method / Results sections with cited specifics.
Results
Section titled “Results”The arXiv abstract is not present in the fetched copy, so quantitative results cannot be cited from the source. Third-party coverage of the launch reports (a) #1 ranking on Artificial Analysis text-to-video Arena at release with Elo ≈ 1269, beating Veo 3, Sora 2, and Runway Gen-4.5 in the same window; (b) explicit “90%+ usable output rate” framing in marketing; (c) frontier-comparable performance on multi-subject interaction and physical-motion scenes. None of these numbers are cited here because the fetched arXiv page does not contain them and the paper page should not pull external claims into the Results section. Treat them as launch-time context; the canonical benchmark column for Seedance 2.0 elsewhere in the wiki is the joint-A+V Open Question that names it as the un-measured closed-source comparison point (Joint audio-video generation Open Questions).
Why it’s interesting
Section titled “Why it’s interesting”This is the long-awaited closed-source datapoint behind the explicit open question on Joint audio-video generation: “Whether closed-source joint A+V systems (Veo 3, Sora 2, Kling 3.0, Seedance 2.0, Vidu Q3) use the same dual-stream-with-bidirectional-cross-attention recipe or something materially different.” The launch-level disclosure (“dual-branch audio-video joint generation”) is consistent with the open consensus established by LTX-2: Efficient Joint Audio-Visual Foundation Model / MOVA: Towards Scalable and Synchronized Video-Audio Generation / SkyReels-V4: Multi-modal Video-Audio Generation, Inpainting and Editing model, but the paper itself is not exposed in the wiki’s fetch — so the question shifts from “is the recipe the same?” to “does the technical report, once accessible, ablate against the open cluster on a shared axis?” Separately, Seedance also appears as a backbone elsewhere in the wiki — FlowAct-R1: Towards Interactive Humanoid Video Generation uses Seedance MMDiT as its base for a streaming AR humanoid video product, so understanding Seedance 2.0’s architecture has downstream implications for what FlowAct-R1’s distillation pipeline is built on top of.
See also
Section titled “See also”- Joint audio-video generation — explicit Open Question now has Seedance 2.0 as a (partially-resolved) closed-source datapoint
- LTX-2: Efficient Joint Audio-Visual Foundation Model — open dual-stream T2AV baseline; LTX-2 is the cleanest open counterpart on capacity-allocation and RoPE alignment
- MOVA: Towards Scalable and Synchronized Video-Audio Generation — second open dual-stream datapoint (MoE, Aligned RoPE)
- SkyReels-V4: Multi-modal Video-Audio Generation, Inpainting and Editing model — third open dual-stream datapoint; symmetric architecture + unified inpainting interface, the closest structural match to Seedance 2.0’s “one model for T2AV + I2V + extension + editing” framing
- FlowAct-R1: Towards Interactive Humanoid Video Generation — downstream product built on top of Seedance MMDiT as the bidirectional teacher
- World Foundation Models — Seedance 2.0 sits in the closed-flagship cohort alongside Genie 3 and the Sora 2 / Veo 3 reference points
- Open foundation-model releases — Seedance 2.0 is the closed counterweight to the open-T2AV cluster on that page
- Dual-stream diffusion transformer — architectural family Seedance 2.0 belongs to per launch disclosure