Pantograph introduces Pan-1 — Minecraft model with RL-based pretraining, arguing video games are the right testbed for robotics
Pantograph (@pantographPBC) announces Pan-1, a Minecraft agent trained with an RL-based pretraining technique the team frames as a potential route to unlocking internet-scale video for robotics foundation models. The launch tweet advertises zero-shot achievement of diverse goals — fighting mobs, building structures, exploring — none of which were specifically trained for. Notable primarily as a product/positioning marker from a robotics-hardware lab (Pantograph builds low-cost RL-driven robot arms) making the explicit argument that video games are a legitimate testbed for embodied methods: reproducible, safe, open-ended, long-horizon.
Key claims
Section titled “Key claims”- Pan-1 is trained with “an RL-based pretraining technique” positioned as generalizing to internet-scale video for robotics models [tweet body].
- The model achieves diverse Minecraft goals — fight mobs, build structures, explore — without being specifically trained on any of them [tweet body].
- Pantograph argues (per the quoted line the Slack sharer endorsed) that video games are a useful testing ground for robotics because they are reproducible, open-ended, safe, and support long-horizon goals — and that “if your method can’t solve video games, it probably won’t solve robotics” [tweet quoted line].
Method
Section titled “Method”Only the launch tweet is available at filing time; no technical report, code, or benchmarks were linked. The framing implies Minecraft is used as the pretraining environment (not just an eval), with reinforcement learning as the training objective. The strategic argument the tweet foregrounds is that Minecraft-scale gameplay video — labeled or otherwise made action-recoverable via RL techniques — is a stepping stone toward exploiting the much larger reservoir of internet gaming video for robotics-oriented foundation models. Details of the “RL-based pretraining technique” are not disclosed in the tweet.
Results
Section titled “Results”Zero-shot diversity of behaviors (combat, construction, exploration) is claimed via a bundled demo video; no benchmark numbers, comparisons to VPT / MineDojo / Voyager / STEVE-1, or ablations are reported in the tweet itself.
Why it’s interesting
Section titled “Why it’s interesting”The tweet is a positioning statement more than a technical release, but it lands in the middle of an active thread on the wiki. It stakes out the RL-pretraining-on-games corner directly adjacent to NitroGen: An Open Foundation Model for Generalist Gaming Agents, which took the behavior-cloning-on-40k-hours-of-controller-overlay-video corner of the same design space — both bet that public gaming video is scalable pretraining substrate for embodied agents, but disagree on the training objective (BC vs RL). It also complements Lumine: An Open Recipe for Building Generalist Agents in 3D Open Worlds (contractor data + VLM-emit-text on Genshin Impact) and SIMA 2: A Generalist Embodied Agent for Virtual Worlds (DeepMind multi-game VLA + Genie-3 self-improvement) as an open, RL-first, Minecraft-anchored fourth corner.
The endorsed argument — video games as a robotics testbed — is a compact version of the position World Foundation Models tracks around video pretraining for embodied AI, and specifically the The flavor of the bitter lesson for computer vision thesis that pixel-in / action-out video-generative pretraining should replace explicit-3D perception stacks. Pantograph’s angle differs: rather than treating video generation as the pretraining objective, they treat game-environment RL as the pretraining objective — a proposition that would compete with, not replace, the video-generative-rollout framing that dominates the WFM cluster.
See also
Section titled “See also”- NitroGen: An Open Foundation Model for Generalist Gaming Agents — sibling foundation model for generalist gaming agents; BC-on-40k-hours-of-controller-overlay-video corner (vs Pan-1’s RL-pretraining corner)
- Lumine: An Open Recipe for Building Generalist Agents in 3D Open Worlds — contractor data + Qwen2-VL emit-text on Genshin Impact; different data + interface corner
- SIMA 2: A Generalist Embodied Agent for Virtual Worlds — DeepMind multi-game VLA + Genie-3 self-improvement; closed-flagship corner
- World Foundation Models — WFM cluster where “video games as robotics testbed” is one live framing
- The flavor of the bitter lesson for computer vision — foundation-layer normative argument for video-generative pretraining as the embodied-AI substrate; Pan-1’s RL-in-games framing is a distinct bet