Jitendra Malik: don't let CV researchers in robotics skip the sensorimotor level
Berkeley CV/robotics professor Jitendra Malik posts unsolicited advice to computer-vision researchers jumping into robotics: don’t over-index on VLMs / VLAs, because the open problems live at the sensorimotor level — manipulation, hand-object interaction, contacts and forces — and proprioception + tactile sensing matter as much as vision. Closes with “you can’t do robotics without doing robotics,” i.e. cherry-picked demos are not evidence of capability. A short opinion tweet (47K views at fetch time), no underlying paper.
Key claims
Section titled “Key claims”- Most open problems in robotics are in manipulation, which centres on hand-object interaction with contacts and forces [tweet body].
- Proprioception and tactile sensing are as important as vision for robotic manipulation [tweet body].
- Cherry-picked demos overstate VLM/VLA capability; doing robotics requires actually doing robotics [tweet body].
Method
Section titled “Method”Not a research artifact — a single-paragraph opinion tweet from a senior CV/robotics researcher (UC Berkeley EECS; Amazon VP & Distinguished Scientist). Filed because it stakes a position on the VLA-vs-sensorimotor split that the wiki has been triangulating across several recent filings.
Results
Section titled “Results”Not applicable. Engagement: 47.3K views, 903 likes, 109 retweets, 302 bookmarks at fetch time.
Why it’s interesting
Section titled “Why it’s interesting”Sets up a useful tension with the VLA line the wiki has been tracking — π*0.6: a VLA That Learns From Experience (RECAP) (π*0.6 / RECAP) and π0.7: A Steerable Robotic Foundation Model with Emergent Compositional Generalization (π0.7) both claim VLAs are the right substrate, with iterated on-policy data + human-gated DAgger covering the contact-rich tasks Malik flags; LingBot-VLA: A Pragmatic VLA Foundation Model (LingBot-VLA) takes the same VLM-backbone-as-policy bet at smaller scale. Malik’s argument is the inverse: the VLA framing under-weights the proprioceptive/tactile side of manipulation. Also relevant context for LEAP: Supercharging LLMs for Formal Mathematics with Agentic Frameworks-style “agents on top of VLMs” debates and for the embodied-agent line in SIMA 2: A Generalist Embodied Agent for Virtual Worlds, where the action substrate is keyboard+mouse rather than sensorimotor.
See also
Section titled “See also”- π*0.6: a VLA That Learns From Experience (RECAP) — the closest counterpoint: a VLA that does contact-rich manipulation by iterating on-policy data + interventions
- π0.7: A Steerable Robotic Foundation Model with Emergent Compositional Generalization — π0.7 as a steerable VLA, explicitly the kind of artifact Malik is warning against over-weighting
- LingBot-VLA: A Pragmatic VLA Foundation Model — LingBot-VLA, another VLM-backboned policy in the line Malik critiques
- SIMA 2: A Generalist Embodied Agent for Virtual Worlds — embodied generalist agent operating at a non-sensorimotor action level