Skip to content

Robots Need More than VLA and World Models

A position paper arguing that the dominant “scale up VLA + world models” framing of generalist robotics is incomplete. The bottleneck is not policy capacity but the lack of mechanisms that convert the world’s abundant unstructured behavioural data — human motion, internet video, simulation rollouts, demonstrations — into grounded robot supervision. The authors name four missing components (“interfaces”) and survey current progress against each: data interfaces (autolabelling), embodiment interfaces (cross-morphology retargeting), world-model interfaces (physics-grounded 3D reasoning), and reward interfaces (task-progress and success inference from video + language).

  • The central bottleneck of generalist robot intelligence is not policy learning but the absence of mechanisms that turn unstructured behavioural data into grounded robot supervision, because most such data lacks embodiment-specific action labels, task semantics, and reward structure [Abstract].
  • Four missing components are required for the next generation of robotics — a data interface for autolabelling unstructured behaviour, an embodiment interface for retargeting human motion to robot actions, a world-model interface for physics-grounded 3D reasoning, and a reward interface for inferring task progress and success from video and language [Abstract].
  • Robot demonstrations alone are insufficient; the proposed research agenda is to learn from “the broader physical world” — human motion, internet video, simulation rollouts, and interactive demonstrations — via the four interfaces rather than via larger VLA models trained on more teleop data [Abstract].

Position paper; no model trained. The paper’s mechanical contribution is a decomposition of the problem of generalist robot intelligence into four interface specifications, each surveyed against current progress in robot foundation models, cross-embodiment datasets, learning from video, world models, and reward modelling. Submitted by Haitham Bou-Ammar’s group (8 co-authors spanning Stanford, ETH Zürich, IIT, TU Darmstadt, Edinburgh) on 4 Jun 2026; v1 only at filing time.

No quantitative results — survey + position paper. The headline claim is qualitative: scaling robot demonstrations and VLA models will not deliver generalist robots without the four named interfaces. The contribution lies in framing rather than benchmarking.

Names a problem the wiki’s robotics-adjacent filings have been circling without articulating directly: every individual piece of the proposed agenda is being worked on in isolation. EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World and Build AI's early egocentric release — 400k action labels, 2.5k clips, 2× open-source dataset size (Eddy Xu tweet, Oct 22 2025) are data-interface plays (egocentric human capture as VLA-ready data); ReActor: Reinforcement Learning for Physics-Aware Motion Retargeting is an embodiment-interface play (bilevel-optimization retargeting of human motion to a robot morphology with an RL tracking policy); Direct Video-Action Models — Causal Video Models Are Data-Efficient Robot Policy Learners and Evaluating Gemini Robotics Policies in a Veo World Simulator are world-model-interface plays (video models as the policy and as the evaluator); and BaseReward: A Strong Baseline for Multimodal Reward Model is the closest filed reward-interface anchor. The paper does not advance any of these axes individually, but it does name a four-way decomposition the wiki has not had as a navigation aid for the robotics cluster — useful primarily as a vocabulary marker rather than as a research result.