Kling Video 2.6 — first Kling AI model with native audio (announcement)
Kling AI’s Day-3 release announces Video 2.6, the first Kling model with native audio generation — joint video+audio in a single model rather than a T2V→V2A pipeline. The launch is paired with a short film “I Have a Secret” co-produced with several Kling Creative Partners (@ViralAiFun, @mad_mask, @JH4TC, @misslaidlaw). No architecture details, parameter counts, or benchmarks are disclosed; the tweet is a product launch with a promo (200 credits for follow+like+RT, 200 Standard-Plan giveaways). The relevance for the team is that Kling 2.6 is now another closed-source datapoint in the joint-audio-video race alongside Veo 3, Sora 2, Seedance 2.0, and Vidu Q3.
Key claims
Section titled “Key claims”- Kling Video 2.6 is Kling AI’s first model with native audio generation, framed as “See the Sound, Hear the Visual” rather than a chained T2V→V2A pipeline [Tweet body].
- The launch demo is a short film “I Have a Secret” co-produced with four Kling Creative Partners [Tweet body].
- No technical disclosure — no parameter count, architecture, training data, audio sample rate, max duration, or benchmark numbers are stated [Tweet body].
Method
Section titled “Method”Not disclosed. The tweet shows a 43-second demo video and lists creative collaborators; there is no accompanying technical report, system card, or model card linked from the tweet. Whether 2.6 uses the dual-stream-DiT-with-bidirectional-cross-attention recipe that the open T2AV cluster has converged on (LTX-2: Efficient Joint Audio-Visual Foundation Model, MOVA: Towards Scalable and Synchronized Video-Audio Generation, SkyReels-V4: Multi-modal Video-Audio Generation, Inpainting and Editing model) or something materially different is exactly the question the Joint audio-video generation concept page already flags as an open one.
Results
Section titled “Results”No quantitative results in the announcement. The 43-second sample is the only public artifact, and there’s no apples-to-apples benchmark comparison against Veo 3 / Sora 2 / open T2AV. The 290.3K views and ~1.5K reposts are engagement metrics, not capability metrics.
Why it’s interesting
Section titled “Why it’s interesting”Kling 2.6 is the third closed-source frontier T2AV model to ship in 2025–2026 after Veo 3 and Sora 2, and the Joint audio-video generation concept page already lists Kling 3.0 as one of the closed-source reference points the open cluster (LTX-2, MOVA, SkyReels-V4) is chasing — so this is the canonical “what is the closed-source recipe?” datapoint, even with no architecture revealed. It complements Kling-MotionControl Technical Report (the only previously-filed Kuaishou Kling artifact, on character motion control) by pinning down that the same group is now also shipping joint audio. The lack of any technical detail is itself informative: contrast with Alibaba’s Wan releases (Wan2.2: World's First Open-Source MoE Video Generation Model (launch announcement), Wan-Image: Pushing the Boundaries of Generative Visual Intelligence) where each version ships an arxiv tech report — Kuaishou is staying closer to the OpenAI/Google “product first, paper maybe later” cadence for their flagship line.
See also
Section titled “See also”- Joint audio-video generation — Kling 2.6 is a new closed-source datapoint in the cluster; the page’s open questions about whether closed-source models use the same dual-stream + bidirectional cross-attention recipe as LTX-2/MOVA/SkyReels-V4 are exactly what this announcement does not answer
- Kling-MotionControl Technical Report — the only other filed Kling artifact, on character motion control; same group (Kuaishou) but the character-animation product line rather than the base T2V/T2AV line
- LTX-2: Efficient Joint Audio-Visual Foundation Model — the strongest open-source comparator and the model the concept page treats as the open T2AV baseline
- Seedance 2.0: Advancing Video Generation for World Complexity — ByteDance’s closed-source video model that also added native audio; comparable launch-grade artifact at comparable secrecy level