Skip to content

AI humans from Runway (Runway Characters / GWM Avatars product announcement)

A Runway tweet, shared as “ai humans from runway.” The tweet itself is age-gated on X, so its exact text and media cannot be read without authentication. Contemporaneous Runway product/news pages describe the Runway Characters product line and the underlying GWM-1 (General World Model) family — specifically the GWM Avatars variant — as the audio-driven, real-time conversational video model that the “ai humans” framing almost certainly refers to. The product surface is fully autonomous, real-time virtual video agents capable of natural conversation, generated from a single reference image with no per-character fine-tuning. Filed as a market-signal pointer, not a technical reference — no paper, no architecture disclosure, no benchmarks.

Note: This is a product announcement, not a paper. Claims below are Runway’s marketing/product statements assembled from the runwayml.com product page and the company’s news/release notes, because the tweet itself is behind an age gate. They are not validated technical results.

  • Runway positions itself as building “foundational General World Models” intended to simulate all possible worlds and experiences, with characters/avatars as one of three product surfaces over the same backbone [Runway product page].
  • GWM-1 is described as Runway’s first General World Model family, built on top of Gen-4.5, autoregressive (frame-by-frame), real-time, and interactively controllable via actions — camera pose, robot commands, audio. It ships in three post-trained variants: GWM Worlds (explorable environments), GWM Avatars (conversational characters), GWM Robotics (manipulation) [Runway news / release notes].
  • Runway Characters is the productized form of GWM Avatars — an audio-driven interactive video generation model that produces fully expressive conversational characters from a single reference image, with the model handling eye movements, lip-sync, facial expressions, and gestures during both speaking and listening, sustained across extended conversations [Runway product page].
  • The product is pitched at use cases like AI customer support, brand IP activation (turning static mascots into interactive personas), and conversational video agents with custom voice / personality / knowledge — i.e. real-time video agents as an API [Runway product page].
  • No technical disclosure: no architecture, no parameter count, no training data, no benchmark numbers, no comparison against open-source audio-conditioned avatar / talking-head systems.

Not disclosed. The closest publicly-stated structural facts are that Runway describes GWM-1 as autoregressive (“generates frame by frame”), real-time, action-conditioned (with audio listed as one of the action modalities), and built on top of their Gen-4.5 video model — so structurally the Avatars variant is a post-trained branch of an AR video DiT with audio + reference-image conditioning. No architecture diagram, no training recipe, no inference-cost numbers, no FPS / resolution specification. No model weights or API documentation is linked from the tweet itself.

None disclosed in the tweet. A separate Runway research artifact (the Gen-4.5 perceptual study, summarized in their release notes) reports that >90% of participants could not reliably distinguish Gen-4.5 outputs from real video in a 20-video forced-choice protocol (1,043 participants, overall detection accuracy 57.1%, only 9.5% achieving statistically significant detection at p<0.05) — this is Gen-4.5 in general, not GWM Avatars specifically, but it’s the closest quantitative signal Runway has released that bears on “AI humans” indistinguishability. No Characters-specific numbers are available.

Adds a Western closed-flagship datapoint to the wiki’s growing map of real-time interactive video model products: structurally parallel to Xmax X1 — Real-Time Interactive Video Model (Product Announcement) (Xmax X1, also closed, also real-time interactive, but pitched at AR camera-feed re-rendering rather than conversational avatars) and to Project Genie: Experimenting with infinite, interactive worlds (Project Genie / Genie 3, also closed, also real-time, but pitched at explorable world simulation). The “single reference image → real-time conversational character” framing is the closest closed-source counterpart to the open-source talking-head / interactive-humanoid lineage filed under FlowAct-R1: Towards Interactive Humanoid Video Generation and DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video Generation — both of which disclose architecture and numbers, neither of which targets real-time conversational use. The interesting new axis vs. Genie 3 / Waymo / Xmax is audio-driven conversational dynamics (eye contact, listening behavior, mid-speech gesture) as the product surface, rather than camera-pose or environment exploration. As with every closed flagship in this cluster, the disclosure asymmetry is the story: open papers reveal how the trick is done; closed products show how good it can look.