Skip to content

Introducing Runway Gen-4.5

Runway announces Gen-4.5, a closed-source text-to-video foundation model that debuts at the top of the Artificial Analysis Text-to-Video Arena with 1,247 Elo points. The post is a product/marketing release: no architecture, no parameter counts, no benchmarks beyond the headline Elo, and no technical disclosure. Runway frames Gen-4.5 simultaneously as a “video generation” model and a “world model” — quoted alongside NVIDIA’s Jensen Huang, who calls it a “video and world model” — and emphasizes pretraining-data efficiency and post-training improvements as the source of the quality gain, while keeping Gen-4’s serving speed and price.

  • Gen-4.5 holds the #1 position on the Artificial Analysis Text-to-Video benchmark at 1,247 Elo, ahead of all other listed models [§Introduction].
  • Gen-4.5’s gains are attributed to “significant advances in both pre-training data efficiency and post-training techniques” rather than to a disclosed architectural change [§Introduction].
  • Gen-4.5 preserves the inference speed and per-plan pricing of Gen-4 — i.e. quality scaling without inference-cost scaling [§Evolution of Video Generation].
  • All Gen-4 control modes (Image-to-Video, Keyframes, Video-to-Video, etc.) are being ported to Gen-4.5 — same product surface, new backbone [§Evolution of Video Generation].
  • Failure modes are explicitly named in the post: causal-reasoning errors (effects preceding causes), object-permanence violations across frames, and a “success bias” toward action outcomes — i.e. acknowledged shortfalls on the same axes RISE-Video and Physion-Eval measure [§marketing footer].
  • Inference is served on NVIDIA Hopper and Blackwell GPUs as part of a partnership covering pretraining, post-training, and inference lifecycle [§High-performance].
  • Runway describes Gen-4.5 as both a video model and a “world model” — Huang’s quote in the post calls it a “video and world model” [§High-performance, quote].

Not disclosed. The post contains no architecture diagram, no parameter count, no training-data sourcing, no objective function, no inference-time scheduler details, no quantitative ablations. Demonstration is by curated example prompts grouped along five qualitative axes — Complex Scenes, Detailed Compositions, Physical Accuracy, Expressive Characters, and a style range covering Photorealistic / Non-photorealistic / Slice of Life / Cinematic. Each axis is shown via three short clips on a fixed prompt. The post is closer in genre to a product page than a technical report; the only numeric claim is the Artificial Analysis Elo (1,247).

  • Headline: 1,247 Elo on Artificial Analysis Text-to-Video Arena, #1 on the leaderboard at announcement [§Introduction].
  • No comparison numbers (relative Elo deltas vs Veo 3 / Sora 2 / Seedance 2.0 / Kling / Hailuo / Wan are not given, though leaderboard position implies an ordering).
  • No quantitative results on any of the physics/reasoning/consistency axes the post itself names as failure modes (no Physion-Eval, no VBench, no PhysicsIQ, no RISE-Video).
  • Qualitative claims: realistic weight/momentum/force, proper liquid dynamics, coherent fine details (hair strands, material weave) across motion and time [§Evolution of Video Generation].

A second public market-signal datapoint from Runway in the wiki’s current ingest window, after Introducing Runway Labs (the Runway Labs incubator) and AI humans from Runway (Runway Characters / GWM Avatars product announcement) (Runway Characters / GWM Avatars) — the three together pin down Runway’s positioning as both a closed-flagship video lab and a self-described world-model lab. Useful as a leaderboard datapoint sitting alongside the wiki’s other top-of-the-stack video releases — Seedance 2.0: Advancing Video Generation for World Complexity (ByteDance Seedance 2.0), Veo 3 Tech Report (Google Veo 3), Waver: Wave Your Way to Lifelike Video Generation (Waver), Moonvalley Marey — controllable AI video model with 360° camera, pose, and motion transfer (Moonvalley Marey) — and as the only one of those whose disclosure is purely product-side with no technical report. Notable that Runway names the same failure modes (causal reasoning, object permanence, action-success bias) that Physion-Eval: Evaluating Physical Realism in Generated Video via Human Reasoning and the RISE-Video benchmark already operationalize quantitatively — but does not measure Gen-4.5 against them.