Skip to content

Kairos-HomeWorld — whole-home 3D scene generation with object-level interactivity (ACE Robotics + CUHK MMLab + Shenzhen Loop Area Institute)

ACE Robotics, CUHK MMLab, and Shenzhen Loop Area Institute announce Kairos-HomeWorld, framed as the first unified framework for whole-home 3D generation with object-level interactivity from a single text prompt. Alongside the model, the team open-sources what is described as the largest whole-home 3D dataset built for Chinese residential environments: 300K real-world floor plans + 5K fully interactive home scenes (press coverage adds 50K physics-enabled interactive object assets and >15 manipulable objects per home, with claims of a four-stage hierarchical pipeline of floorplan generation → 2D-to-3D lifting → recursive refinement → manipulable-object placement). Pitched as simulation infrastructure for embodied AI / humanoid robot training, explicitly addressing the limitation that prior indoor-scene generation has been restricted to single rooms with weak global consistency.

  • Generates complete residential environments that are “globally consistent, physics-compliant, and simulation-ready” from a single text prompt [tweet OP].
  • Open-source dataset: 300K real-world Chinese residential floor plans + 5K fully interactive home scenes [tweet OP].
  • Positioned as infrastructure for embodied AI as companies like Figure AI expand robot training into residential environments [tweet OP].

The tweet itself does not describe architecture. Synchronous press coverage (Media OutReach Newswire, June 5 2026) attributes a four-stage hierarchical architecture: (1) floorplan generation, (2) 2D-to-3D lifting, (3) recursive refinement, (4) manipulable-object placement. Floor plans are vectorized from real-world listings via a multi-stage automated pipeline that labels door/window positions, room geometry, functional zoning, and connectivity. The 5K furnished homes average >15 physics-enabled manipulable objects per home, powered by the “PhysX-Omni” object-asset model. No paper, code, or architecture document is linked from the tweet at filing time — claims are tracked here from the announcement only.

No quantitative benchmark numbers are released in the tweet. Press coverage cites a “Footprint Object Density of 4.16, the highest among compared methods,” with no leaderboard, ablations, or baseline list disclosed. RPLAN (~80K floor plans) and ResPlan (~17K) are referenced as the closest dataset comparisons — Kairos-HomeWorld’s 300K is ~4× the larger of the two. The company asserts deployment “in ACE ROBOTICS’ daily robot training” already.

Anchors a recurring 2026 thread on the wiki: text-prompted whole-scene generation with explicit-state outputs marketed at embodied AI. Most directly comparable to SAGE: Scalable Agentic 3D Scene Generation for Embodied AI (agentic 3D scene generation from a task description, embodied-AI framing) and WorldGen: From Text to Traversable and Interactive 3D Worlds (Meta Reality Labs’ text → traversable mesh-based 3D worlds), but extends both with a whole-home (not single-room, not generic-scene) framing and a purpose-built localized dataset (Chinese residential floor plans). Complements The Role of Simulation in Scalable Robotics, Genesis World 1.0, and the Path Forward on the bet that explicit-physics + scalable simulation infrastructure is still the trustworthy substrate for robot training — counterweight to the neural-rollout WFM line (PAN, Genie 3). Until a paper or code drop, the announcement should be treated as a product/press marker rather than a technical primitive.