Skip to content

Lerrel Pinto — in-context learning for robots is not that hard

Lerrel Pinto (NYU robotics) tweets a short-form assertion with an attached video demo: “Turns out that doing In-context learning for robots is not that hard.” No paper or blog is linked; the tweet is the artifact. It joins a fast-growing 2026 chorus — Skild S1, Generalist GEN-1.5, HOST, RoboTTT — that a video demonstration prompt at inference time can drive competent manipulation without gradient updates, and hints that the recipe is less exotic than concurrent write-ups suggest.

  • ICL for robot manipulation policies is achievable without a specialized meta-learning outer loop or exotic architecture — implied by the “not that hard” framing and demo video [tweet body].

None disclosed. The tweet is an assertion + a short attached video showing a robot executing a task. No architecture, training data, prompt format, task, or evaluation protocol is described. Follow-up threads and replies (visible in the fetched page — @notismaelvega, @breadli428, @kamalgupta09) are reactions, not method details. Reply from @breadli428 raises the standard critique that “OOD-ness” of the tested task is not well-defined, echoing Anirudha Majumdar — pure zero-shot task inference is the stronger test of robot foundation models‘s methodology point.

None reported. A single video demo is shown in the post; no success rate, task list, comparison, or scaling curve.

Data-point on how the ICL-for-robots consensus is forming: four in-scope entries in the last month (S1: In-Context Learning for Robotics, GEN-1.5: Embodied Foundation Models are One-Shot Learners, Robots Acquire Manipulation Skills in Seconds from a Single Human Video (HOST), plus this NYU tweet) all making the same qualitative claim from different labs, with Skild’s S1 supplying the only filed matched-recipe scaling law. Pinto’s “not that hard” is the shortest possible endorsement of the S1 recipe framing — that ICL emerges from ordinary demonstration-conditioned pretraining without an ICL-specific architecture — while carrying zero verifiable evidence. Load-bearing caveat for future readers: the tweet is a claim, not a paper; the sibling reply from @breadli428 correctly names the missing evaluation methodology (OOD-ness definition), which is exactly the gap Anirudha Majumdar — pure zero-shot task inference is the stronger test of robot foundation models and Skild’s L2–L5 shift ladder in S1: In-Context Learning for Robotics have started to fill.