Wan 2.5 live on WaveSpeed — 5s/10s clips up to 1080p (AI Pulse correction tweet)
A short correction tweet from AI Pulse (@youraipulse) on 23 Sep 2025 noting that Wan 2.5 — which AI Pulse had previously teased as “coming tomorrow” via an Alibaba 2.5-Preview leak — was already serving via the WaveSpeed AI API: 5-second and 10-second clips, up to 1080p. The tweet itself is a one-line product-availability nudge; the underlying release (Wan 2.5) is Alibaba’s natively multimodal text/image/video/audio model with one-pass A/V synchronization, positioned as an open-API competitor to Veo 3.
Key claims
Section titled “Key claims”- Wan 2.5 is reachable via WaveSpeed AI’s API at the time of the tweet, supporting 5s and 10s generation durations and up to 1080p output [tweet body].
- AI Pulse’s prior tweet (quote-RT’d here) had teased Wan 2.5-Preview as “coming tomorrow” based on an Alibaba teaser image — this tweet is the correction that it’s already live [tweet body, quote-RT].
Method
Section titled “Method”Product-availability tweet, no methodology. The substantive content is in the WaveSpeed product page and accompanying launch posts (linked under “See also”), which describe the underlying Wan 2.5 model: a unified text/image/video/audio backbone with one-pass A/V synchronization (lip-sync + sound effects + background audio), 480p / 720p / 1080p at 24fps, durations up to 10s, six aspect-ratio options, and multilingual prompting (English + Chinese particularly strong).
Results
Section titled “Results”Not applicable — this is a tweet announcing API availability, not a release with measured results. WaveSpeed’s own marketing materials position Wan 2.5 as “up to 3× cheaper than Google Veo 3” with comparable A/V-sync quality, but no third-party benchmarks are cited in the tweet itself.
Why it’s interesting
Section titled “Why it’s interesting”Wan 2.5 is the version of the Alibaba Wan video stack that ships native one-pass audio-video generation — placing it in the same architectural slot as the joint-A/V work Luma already tracks. It complements the open-weights UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation (UniForm — unified A/V diffusion transformer) and Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation (Ovi — twin-backbone cross-modal fusion) by being the closed-API datapoint for the same capability at frontier scale, and contrasts with UniVerse-1: Unified Audio-Video Generation via Stitching of Experts (UniVerse-1 — A/V via stitching of separately-trained experts) on the unified-vs-stitched axis. On the release-trajectory side, it slots between Wan2.2: World's First Open-Source MoE Video Generation Model (launch announcement) (Wan2.2 — MoE video, no audio) and Wan 2.7 upcoming features preview (video editing, V2V, real-time) (Wan 2.7 preview — V2V editing, real-time) in the Wan family’s continued-development lineage; the fact that 2.5 ships first as a WaveSpeed-hosted API rather than an open-weights drop (unlike 2.2) is itself a signal worth tracking for Open foundation-model releases.
See also
Section titled “See also”- Wan2.2: World's First Open-Source MoE Video Generation Model (launch announcement) — Wan2.2 launch (the open-weights MoE predecessor)
- Wan 2.7 upcoming features preview (video editing, V2V, real-time) — Wan 2.7 preview (next-version downstream successor)
- UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation — open-weights single-backbone joint A/V diffusion, contemporary architectural counterpart
- Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation — twin-backbone cross-modal fusion for A/V, alternative architecture
- UniVerse-1: Unified Audio-Video Generation via Stitching of Experts — A/V via stitching of experts (contrast with Wan 2.5’s claimed unified backbone)
- Joint audio-video generation — broader concept cluster
- Open foundation-model releases — Wan 2.5 is API-only at this point; tracked as a release-trajectory datapoint