Lerrel Pinto — in-context learning for robots is not that hard
Lerrel Pinto (NYU robotics) tweets a short-form assertion with an attached video demo: “Turns out that doing In-context learning for robots is not that hard.” No paper or blog is linked; the tweet is the artifact. It joins a fast-growing 2026 chorus — Skild S1, Generalist GEN-1.5, HOST, RoboTTT — that a video demonstration prompt at inference time can drive competent manipulation without gradient updates, and hints that the recipe is less exotic than concurrent write-ups suggest.
Key claims
Section titled “Key claims”- ICL for robot manipulation policies is achievable without a specialized meta-learning outer loop or exotic architecture — implied by the “not that hard” framing and demo video [tweet body].
Method
Section titled “Method”None disclosed. The tweet is an assertion + a short attached video showing a robot executing a task. No architecture, training data, prompt format, task, or evaluation protocol is described. Follow-up threads and replies (visible in the fetched page — @notismaelvega, @breadli428, @kamalgupta09) are reactions, not method details. Reply from @breadli428 raises the standard critique that “OOD-ness” of the tested task is not well-defined, echoing Anirudha Majumdar — pure zero-shot task inference is the stronger test of robot foundation models‘s methodology point.
Results
Section titled “Results”None reported. A single video demo is shown in the post; no success rate, task list, comparison, or scaling curve.
Why it’s interesting
Section titled “Why it’s interesting”Data-point on how the ICL-for-robots consensus is forming: four in-scope entries in the last month (S1: In-Context Learning for Robotics, GEN-1.5: Embodied Foundation Models are One-Shot Learners, Robots Acquire Manipulation Skills in Seconds from a Single Human Video (HOST), plus this NYU tweet) all making the same qualitative claim from different labs, with Skild’s S1 supplying the only filed matched-recipe scaling law. Pinto’s “not that hard” is the shortest possible endorsement of the S1 recipe framing — that ICL emerges from ordinary demonstration-conditioned pretraining without an ICL-specific architecture — while carrying zero verifiable evidence. Load-bearing caveat for future readers: the tweet is a claim, not a paper; the sibling reply from @breadli428 correctly names the missing evaluation methodology (OOD-ness definition), which is exactly the gap Anirudha Majumdar — pure zero-shot task inference is the stronger test of robot foundation models and Skild’s L2–L5 shift ladder in S1: In-Context Learning for Robotics have started to fill.
See also
Section titled “See also”- S1: In-Context Learning for Robotics — Skild’s full-writeup version of the same claim, with matched-recipe 1k→100k-hour scaling curve
- GEN-1.5: Embodied Foundation Models are One-Shot Learners — concurrent one-shot ICL claim (short-horizon in-distribution)
- Robots Acquire Manipulation Skills in Seconds from a Single Human Video (HOST) — HOST: explicit-cascade one-shot recipe with published numbers
- Anirudha Majumdar — pure zero-shot task inference is the stronger test of robot foundation models — methodology counter: strip the prompt, measure pure zero-shot task inference
- VLA Models — adds another public claim on the ICL-native VLA recipe row