Skip to content

HappyHorse 1.0: #1 State-of-the-Art Joint Audio-Video Generator (Alibaba Taotian Future Life Lab)

HappyHorse 1.0 is a pseudonymous video+audio generative model that topped the Artificial Analysis Video Arena in early April 2026, eventually claimed by a team at Alibaba’s Taotian Future Life Lab (led by Zhang Di, former VP of Kuaishou and technical lead of Kling). The landing page advertises a unified single-pass joint audio-video architecture and promises “everything open” (base + distilled + super-resolution + inference code), but as of filing the GitHub and Model Hub links are still “coming soon” and fal’s official partner page explicitly states the model will be closed source — putting the open-source claim in doubt. The relevance to Luma is the Arena result itself (claimed #1 in T2V/I2V without audio, ahead of Dreamina Seedance 2.0 by ~60 Elo points in T2V) plus the architectural claim of joint A+V generation, which would slot it next to LTX-2, MOVA, and SkyReels-V4 if details ever ship.

  • HappyHorse-1.0 is a joint video+audio generative model that produces synchronized video and audio from one text prompt in multiple resolutions and languages [landing page §Overview].
  • The team claims a #1 ranking on the Artificial Analysis Video Arena for text-to-video and image-to-video without audio, and #2 with audio (behind Seedance 2.0) [landing page §Rankings, citing the Artificial Analysis April 7, 2026 tweet embedded on the page].
  • The model is built by the Taotian Future Life Lab team at Alibaba (ATH division), led by Zhang Di, previously described as VP at Kuaishou and technical lead of Kling AI [landing page §Team; later confirmed by CNBC reporting on April 10, 2026].
  • The landing page claims a full open release is forthcoming covering “base model, distilled model, super-resolution model, and inference code” with GitHub and Model Hub links marked “coming soon” [landing page §Open Source].
  • The landing page makes no public claims about parameter count, training data, or inference latency — all such numbers (“15B parameters, 40-layer unified self-attention Transformer, joint video+audio in one forward pass, ~38 s for 1080p on a single H100”) are from third-party community compilations and have not been verified [external sources only].

The landing page is a marketing site rather than a technical report; no architecture diagram, training recipe, or benchmark methodology is disclosed. The headline framing is “one model generates video and audio together” in a single forward pass (i.e. not a cascaded T2V→V2A pipeline), with full technical details promised at open-source release. The page links to the Artificial Analysis Video Arena and embeds tweets from Artificial Analysis (April 7, 2026) and several community accounts showing comparison clips against Seedance 2.0.

  • Claimed #1 on Artificial Analysis Text-to-Video and Image-to-Video (no audio) leaderboards as of early April 2026 [landing page rankings, citing Artificial Analysis tweet].
  • Claimed #2 on Text-to-Video and Image-to-Video with audio (behind Seedance 2.0), with the gap narrowing to ~14 Elo points in T2V-with-audio and ~1 Elo point in I2V-with-audio per third-party reporting.
  • No internal benchmark, FVD, or human-preference number is given on the landing page itself — the entire results section is an Arena leaderboard claim plus embedded comparison clips.

If the joint-video+audio claim holds up at open release, HappyHorse-1.0 becomes the fourth datapoint in the open dual-stream joint-A+V cluster filed in Joint audio-video generation (after LTX-2: Efficient Joint Audio-Visual Foundation Model, MOVA: Towards Scalable and Synchronized Video-Audio Generation, and SkyReels-V4: Multi-modal Video-Audio Generation, Inpainting and Editing model), and would be the first to actually top the Artificial Analysis Arena — SkyReels-V4 also reports #1 but on a different snapshot, and the head-to-head against Seedance 2.0 is the closest open vs closed comparison the Arena has produced. Two things make this entry asymmetrically important to track despite the absence of a technical report: (a) the open-source claim is in direct conflict with the fal partner page’s “closed source, not licensable” statement, exactly the kind of marketing-vs-reality gap Open foundation-model releases needs to track to keep “open” meaningful; (b) the unified-single-pass framing, if true at ~15B parameters with no cross-attention, would be a third architectural family next to the dual-stream DiT consensus and the discrete tri-modal MDM of The Design Space of Tri-Modal Masked Diffusion Models. Until weights ship, this is a leaderboard datapoint with a stealth-drop pedigree (cf. Pony Alpha / GLM-5), not an evaluable artifact.