Skip to content

Scaling Robotic Manipulation via Structured World Models and Tactile Sensing — Yunzhu Li (Montreal Robotics)

Invited talk by Yunzhu Li (Columbia, MIT PhD) at the Montreal Robotics seminar series, also delivered at Princeton (March 2026) under the same title. The thesis: scaling robotic manipulation requires both predictive world models that capture how the world evolves under action and rich sensing of physical contact — vision alone is insufficient. The talk argues structured world models (action-conditioned predictive models with explicit physics structure, as opposed to monolithic pixel-rollout video models) paired with tactile sensing are the missing ingredients for closing the manipulation gap, and frames this as the bottleneck distinguishing today’s locomotion-heavy humanoid demos from reliable in-the-wild manipulation.

  • Vision-only perception is the ceiling on current manipulation systems; rich physical-contact sensing is the missing input modality required to scale beyond demo-grade success rates [talk abstract].
  • Structured world models — predictive models of how the world evolves under action, with explicit physical structure — are the predictive substrate manipulation policies need at scale [talk abstract, title].
  • Tactile sensing is positioned as a co-equal component to world modeling, not as an add-on; the talk pairs the two as the joint axis along which manipulation scales [talk title; tweet summary].

(No transcript was available at filing time, so claims here are anchored to the talk title, the publicly posted abstract, and the framing in the tweet pointing at the talk. Section-level anchors will be added in a /bud refresh once a transcript or slides are available.)

The Montreal Robotics video and the Princeton seminar listing both give the same one-line abstract and title — “Scaling Robotic Manipulation via Structured World Models and Tactile Sensing” — without further public detail. Based on Li’s recent publication record at Columbia (structured world models for manipulation, action-conditioned scene graphs, tactile sensing, ManipulationNet benchmark tracks), the talk likely covers (a) action-conditioned predictive models with explicit physical structure as the manipulation backbone, (b) tactile sensors and contact representations as the second input modality, and (c) integration of foundation models as structural priors via VLM-for-task-interpretation + constrained-optimization-for-motion-planning splits. None of this is claimed by the filed video page itself; transcript pending.

No quantitative results are surfaced in the publicly available video metadata or seminar listings; the talk is a position/synthesis lecture rather than a paper-anchored result-presentation. A refresh once a transcript or slide deck appears will pull in the actual numbers Li discusses.

This is the wiki’s first filed invited-talk artifact taking the position that structured world models + tactile sensing are the joint axis for scaling manipulation — sharper and narrower than the broader generative-rollout vs predictive-latent debate the World Foundation Models cluster currently hosts. It complements A Functional Taxonomy of World Models‘s three-way renderer / simulator / planner split by adding a fourth orthogonal axis: what the policy senses, where tactile/contact data is the part vision-anchored WFMs systematically miss. It contrasts with The flavor of the bitter lesson for computer vision‘s “explicit 3D should dissolve into video pre-training” position — Li’s framing keeps explicit physical structure (and explicit non-visual sensing) as first-class, not as artifacts to be absorbed. Directly relevant to the manipulation-evaluation question raised by Dream.exe: Can Video Generation Models Dream Executable Robot Manipulation? (can video models generate executable manipulation?) — Li’s answer would be “not without contact sensing in the loop”.