Skip to content

Jitendra Malik: don't let CV researchers in robotics skip the sensorimotor level

Berkeley CV/robotics professor Jitendra Malik posts unsolicited advice to computer-vision researchers jumping into robotics: don’t over-index on VLMs / VLAs, because the open problems live at the sensorimotor level — manipulation, hand-object interaction, contacts and forces — and proprioception + tactile sensing matter as much as vision. Closes with “you can’t do robotics without doing robotics,” i.e. cherry-picked demos are not evidence of capability. A short opinion tweet (47K views at fetch time), no underlying paper.

  • Most open problems in robotics are in manipulation, which centres on hand-object interaction with contacts and forces [tweet body].
  • Proprioception and tactile sensing are as important as vision for robotic manipulation [tweet body].
  • Cherry-picked demos overstate VLM/VLA capability; doing robotics requires actually doing robotics [tweet body].

Not a research artifact — a single-paragraph opinion tweet from a senior CV/robotics researcher (UC Berkeley EECS; Amazon VP & Distinguished Scientist). Filed because it stakes a position on the VLA-vs-sensorimotor split that the wiki has been triangulating across several recent filings.

Not applicable. Engagement: 47.3K views, 903 likes, 109 retweets, 302 bookmarks at fetch time.

Sets up a useful tension with the VLA line the wiki has been tracking — π*0.6: a VLA That Learns From Experience (RECAP) (π*0.6 / RECAP) and π0.7: A Steerable Robotic Foundation Model with Emergent Compositional Generalization (π0.7) both claim VLAs are the right substrate, with iterated on-policy data + human-gated DAgger covering the contact-rich tasks Malik flags; LingBot-VLA: A Pragmatic VLA Foundation Model (LingBot-VLA) takes the same VLM-backbone-as-policy bet at smaller scale. Malik’s argument is the inverse: the VLA framing under-weights the proprioceptive/tactile side of manipulation. Also relevant context for LEAP: Supercharging LLMs for Formal Mathematics with Agentic Frameworks-style “agents on top of VLMs” debates and for the embodied-agent line in SIMA 2: A Generalist Embodied Agent for Virtual Worlds, where the action substrate is keyboard+mouse rather than sensorimotor.