Scaling Robotic Manipulation via Structured World Models and Tactile Sensing — Yunzhu Li (Montreal Robotics)
Invited talk by Yunzhu Li (Columbia, MIT PhD) at the Montreal Robotics seminar series, also delivered at Princeton (March 2026) under the same title. The thesis: scaling robotic manipulation requires both predictive world models that capture how the world evolves under action and rich sensing of physical contact — vision alone is insufficient. The talk argues structured world models (action-conditioned predictive models with explicit physics structure, as opposed to monolithic pixel-rollout video models) paired with tactile sensing are the missing ingredients for closing the manipulation gap, and frames this as the bottleneck distinguishing today’s locomotion-heavy humanoid demos from reliable in-the-wild manipulation.
Key claims
Section titled “Key claims”- Vision-only perception is the ceiling on current manipulation systems; rich physical-contact sensing is the missing input modality required to scale beyond demo-grade success rates [talk abstract].
- Structured world models — predictive models of how the world evolves under action, with explicit physical structure — are the predictive substrate manipulation policies need at scale [talk abstract, title].
- Tactile sensing is positioned as a co-equal component to world modeling, not as an add-on; the talk pairs the two as the joint axis along which manipulation scales [talk title; tweet summary].
(No transcript was available at filing time, so claims here are anchored to the talk title, the publicly posted abstract, and the framing in the tweet pointing at the talk. Section-level anchors will be added in a /bud refresh once a transcript or slides are available.)
Method
Section titled “Method”The Montreal Robotics video and the Princeton seminar listing both give the same one-line abstract and title — “Scaling Robotic Manipulation via Structured World Models and Tactile Sensing” — without further public detail. Based on Li’s recent publication record at Columbia (structured world models for manipulation, action-conditioned scene graphs, tactile sensing, ManipulationNet benchmark tracks), the talk likely covers (a) action-conditioned predictive models with explicit physical structure as the manipulation backbone, (b) tactile sensors and contact representations as the second input modality, and (c) integration of foundation models as structural priors via VLM-for-task-interpretation + constrained-optimization-for-motion-planning splits. None of this is claimed by the filed video page itself; transcript pending.
Results
Section titled “Results”No quantitative results are surfaced in the publicly available video metadata or seminar listings; the talk is a position/synthesis lecture rather than a paper-anchored result-presentation. A refresh once a transcript or slide deck appears will pull in the actual numbers Li discusses.
Why it’s interesting
Section titled “Why it’s interesting”This is the wiki’s first filed invited-talk artifact taking the position that structured world models + tactile sensing are the joint axis for scaling manipulation — sharper and narrower than the broader generative-rollout vs predictive-latent debate the World Foundation Models cluster currently hosts. It complements A Functional Taxonomy of World Models‘s three-way renderer / simulator / planner split by adding a fourth orthogonal axis: what the policy senses, where tactile/contact data is the part vision-anchored WFMs systematically miss. It contrasts with The flavor of the bitter lesson for computer vision‘s “explicit 3D should dissolve into video pre-training” position — Li’s framing keeps explicit physical structure (and explicit non-visual sensing) as first-class, not as artifacts to be absorbed. Directly relevant to the manipulation-evaluation question raised by Dream.exe: Can Video Generation Models Dream Executable Robot Manipulation? (can video models generate executable manipulation?) — Li’s answer would be “not without contact sensing in the loop”.
See also
Section titled “See also”- World Foundation Models — talk’s framing of “predictive models of how the world evolves under action” sits squarely in this cluster, on the structured/predictive side rather than the generative-rollout side
- A Functional Taxonomy of World Models — companion taxonomic essay; both file in the same week and frame what a world model is for, but Li adds the sensing-modality dimension Fei-Fei’s renderer/simulator/planner split doesn’t cover
- The flavor of the bitter lesson for computer vision — clean counter-position; Sitzmann argues explicit structure should dissolve, Li argues structure + tactile is the way to scale
- Dream.exe: Can Video Generation Models Dream Executable Robot Manipulation? — Dream.exe operationalizes “video model → executable manipulation” and finds visual quality doesn’t predict executability; consistent with Li’s argument that contact sensing is the missing axis
- The Role of Simulation in Scalable Robotics, Genesis World 1.0, and the Path Forward — Genesis World’s “simulation is for evaluation, with explicit physics + path-traced rendering” argument is the simulation-side analog of Li’s structured-world-models position