Skip to content

Project Genie: Experimenting with infinite, interactive worlds

Google DeepMind has rolled out Project Genie — a Google Labs web prototype powered by the Genie 3 world model — to Google AI Ultra subscribers in the US. The system lets a user describe a world via text and images (assisted by Nano Banana Pro for “World Sketching”), step into it, navigate it in real time at 20–24 fps / 720p, and remix existing worlds. It is positioned as the productization step after the August 2025 Genie 3 research preview, and as a step toward AGI agents that can reason in the diversity of the real world. The post is product framing rather than a technical report: no architecture details, no benchmarks, no comparisons to prior systems.

  • Project Genie is rolling out to Google AI Ultra subscribers (US, 18+) starting Jan 29, 2026, as an experimental Google Labs prototype [§rollout].
  • The prototype is powered by Genie 3, presented as a general-purpose world model that generates “the path ahead in real time” as the user moves and interacts, rather than rendering a static 3D snapshot [§Advancing world models].
  • The product exposes three capabilities: World Sketching (text + uploaded/generated images via Nano Banana Pro produce a preview image and let the user define perspective and character), World Exploration (real-time navigable environment with adjustable camera), and World Remixing (build on top of others’ prompts, browse a curated gallery, download videos of explorations) [§How Project Genie works].
  • Genie 3 is described as simulating physics and interactions for dynamic worlds, with “breakthrough consistency” claimed across long-form scenarios spanning robotics, animation, fiction, and historical settings [§Advancing world models].
  • Stated current limitations: outputs may not be true-to-life, may not adhere closely to prompts/images or to real-world physics; character controllability is inconsistent and can show higher latency; generations are capped at 60 seconds [§Building responsibly].
  • Promptable events (announced for Genie 3 in August 2025) are explicitly not yet exposed in this prototype [§Building responsibly].
  • Framed as supporting Google DeepMind’s AGI mission — world models are positioned as the ingredient that lets agents go beyond narrow-domain mastery (Chess, Go) toward “the diversity of the real world” [§Advancing world models].

Not disclosed. The post is a product announcement; the underlying Genie 3 model is described qualitatively (general-purpose, real-time path generation, physics + interaction simulation, multi-modal text+image conditioning, integrated with Nano Banana Pro for image preview and Gemini for unspecified roles). External coverage attributes an 11B-parameter autoregressive transformer running at 720p / 24 fps to Genie 3, but those numbers are not in the blog post itself.

The user-facing flow is: prompt the world (text + images), choose locomotion mode (walking, riding, flying, driving, …) and viewpoint (first/third person), define a character, get a World Sketching preview from Nano Banana Pro, then enter the world and navigate it. Worlds are bounded at 60s per generation. A remix path lets users build on top of prompts in a curated gallery.

No quantitative results. The post offers no benchmarks, no ablations, no comparisons to prior world models (DeepMind’s own Genie 2, NVIDIA Cosmos, Decart Mirage, open-source LingBot-World, etc.). It claims qualitative “breakthrough consistency” relative to static-snapshot exploration but does not substantiate the claim. The only concrete operating parameters mentioned are the 60-second generation cap and the eligibility envelope (US Google AI Ultra subscribers, 18+).

This is the closed-source/proprietary counterpart to the open Genie-3-class systems the wiki is already tracking — most directly Advancing Open-source World Models (LingBot-World) (LingBot-World), which explicitly positions itself against Genie 3 on the open/closed axis, and the NVIDIA Cosmos stack in World Foundation Models. For Luma the interesting bit is the product surface: World Sketching (image-grounded preview pre-commit), World Remixing (a gallery + remix layer that turns single-shot generations into a social object), and the locomotion/perspective conditioning UI. These are interface choices for a world model rather than model choices, and they suggest where the medium is heading even though the post discloses nothing about the model itself. The explicit not-yet-shipped feature — promptable events that mutate the world mid-exploration — is the obvious next axis to watch and overlaps with the action-conditioned-rollout work tracked under Autoregressive Video Generation (FlowAct-R1, Context Forcing, LongVie 2, LingBot-World).