Project Genie: Experimenting with infinite, interactive worlds
Google DeepMind has rolled out Project Genie — a Google Labs web prototype powered by the Genie 3 world model — to Google AI Ultra subscribers in the US. The system lets a user describe a world via text and images (assisted by Nano Banana Pro for “World Sketching”), step into it, navigate it in real time at 20–24 fps / 720p, and remix existing worlds. It is positioned as the productization step after the August 2025 Genie 3 research preview, and as a step toward AGI agents that can reason in the diversity of the real world. The post is product framing rather than a technical report: no architecture details, no benchmarks, no comparisons to prior systems.
Key claims
Section titled “Key claims”- Project Genie is rolling out to Google AI Ultra subscribers (US, 18+) starting Jan 29, 2026, as an experimental Google Labs prototype [§rollout].
- The prototype is powered by Genie 3, presented as a general-purpose world model that generates “the path ahead in real time” as the user moves and interacts, rather than rendering a static 3D snapshot [§Advancing world models].
- The product exposes three capabilities: World Sketching (text + uploaded/generated images via Nano Banana Pro produce a preview image and let the user define perspective and character), World Exploration (real-time navigable environment with adjustable camera), and World Remixing (build on top of others’ prompts, browse a curated gallery, download videos of explorations) [§How Project Genie works].
- Genie 3 is described as simulating physics and interactions for dynamic worlds, with “breakthrough consistency” claimed across long-form scenarios spanning robotics, animation, fiction, and historical settings [§Advancing world models].
- Stated current limitations: outputs may not be true-to-life, may not adhere closely to prompts/images or to real-world physics; character controllability is inconsistent and can show higher latency; generations are capped at 60 seconds [§Building responsibly].
- Promptable events (announced for Genie 3 in August 2025) are explicitly not yet exposed in this prototype [§Building responsibly].
- Framed as supporting Google DeepMind’s AGI mission — world models are positioned as the ingredient that lets agents go beyond narrow-domain mastery (Chess, Go) toward “the diversity of the real world” [§Advancing world models].
Method
Section titled “Method”Not disclosed. The post is a product announcement; the underlying Genie 3 model is described qualitatively (general-purpose, real-time path generation, physics + interaction simulation, multi-modal text+image conditioning, integrated with Nano Banana Pro for image preview and Gemini for unspecified roles). External coverage attributes an 11B-parameter autoregressive transformer running at 720p / 24 fps to Genie 3, but those numbers are not in the blog post itself.
The user-facing flow is: prompt the world (text + images), choose locomotion mode (walking, riding, flying, driving, …) and viewpoint (first/third person), define a character, get a World Sketching preview from Nano Banana Pro, then enter the world and navigate it. Worlds are bounded at 60s per generation. A remix path lets users build on top of prompts in a curated gallery.
Results
Section titled “Results”No quantitative results. The post offers no benchmarks, no ablations, no comparisons to prior world models (DeepMind’s own Genie 2, NVIDIA Cosmos, Decart Mirage, open-source LingBot-World, etc.). It claims qualitative “breakthrough consistency” relative to static-snapshot exploration but does not substantiate the claim. The only concrete operating parameters mentioned are the 60-second generation cap and the eligibility envelope (US Google AI Ultra subscribers, 18+).
Why it’s interesting
Section titled “Why it’s interesting”This is the closed-source/proprietary counterpart to the open Genie-3-class systems the wiki is already tracking — most directly Advancing Open-source World Models (LingBot-World) (LingBot-World), which explicitly positions itself against Genie 3 on the open/closed axis, and the NVIDIA Cosmos stack in World Foundation Models. For Luma the interesting bit is the product surface: World Sketching (image-grounded preview pre-commit), World Remixing (a gallery + remix layer that turns single-shot generations into a social object), and the locomotion/perspective conditioning UI. These are interface choices for a world model rather than model choices, and they suggest where the medium is heading even though the post discloses nothing about the model itself. The explicit not-yet-shipped feature — promptable events that mutate the world mid-exploration — is the obvious next axis to watch and overlaps with the action-conditioned-rollout work tracked under Autoregressive Video Generation (FlowAct-R1, Context Forcing, LongVie 2, LingBot-World).
See also
Section titled “See also”- World Foundation Models — Project Genie is the closed-source flagship in this cluster alongside NVIDIA Cosmos; the LingBot-World entry explicitly benchmarks itself against Genie 3
- Autoregressive Video Generation — real-time path-ahead generation as the user moves is the same long-horizon AR-video problem this concept tracks, just productized
- Advancing Open-source World Models (LingBot-World) — LingBot-World: the closest open-source datapoint to a Genie-3-class interactive world model; same Jan 29, 2026 timing is coincidence worth noting
- NVIDIA Unveils New Open Models, Data and Tools to Advance AI Across Every Industry — NVIDIA’s CES 2026 Cosmos stack release as the other closed/semi-open WFM bundle from the same window