Skip to content

Why Robotics Still Isn't Solved - But Could Be Soon | YC Paper Club

A YC Paper Club episode framed as a roadmap for the remaining obstacles to generalist robotics, structured around four named roadblocks (the sim-to-real gap, action representation, the sensorimotor problem, embodiment drift) and five recipe-side themes (giving robot policies memory, teaching models what’s worth reasoning about, dexterous tool use learned entirely in simulation, why the next great robotics companies will start with teleoperation, running world-action models without two GB200s per robot). Aimed at a startup-founder audience rather than a research audience — indexes the current recipe-lever debate rather than staking a new position. No transcript was retrievable at filing time; the framing is reconstructed from the Slack description and the companion post at ycrootaccess.com.

  • The episode enumerates four remaining roadblocks: the sim-to-real gap, action representation, the sensorimotor problem, and embodiment drift [description].
  • Five follow-on themes are covered: policy memory, what’s worth reasoning about, sim-only dexterous tool use, teleoperation as a startup wedge, and running world-action models without two GB200s per robot [description].

Panel-format paper-club discussion; no accompanying paper, benchmark, or code release. The video itself is the artifact — a taxonomy of open problems and current recipes framed for a founder audience. The written companion (linked from the promotional page) collects transcript and reading list.

No quantitative results are reported. The value is the framing: an outside-the-lab enumeration of what practitioners consider the load-bearing open problems in 2026 robotics, useful as a cross-check against the recipe-lever board that VLA Models tracks internally.

Every one of the four named roadblocks maps cleanly onto positions already on the VLA Models concept page: the sensorimotor problem is Jitendra Malik’s counter-position pushed by Jitendra Malik: don't let CV researchers in robotics skip the sensorimotor level and VLA에게 부족한 결정적 감각 | 로봇에게 Force·Tactile이 반드시 필요한 이유 (Why VLAs Need Force and Tactile Sensing for Robot Manipulation), with concrete recipe answers now filed (FACT, CHORD, Tactile-Reactive Dexterous Hand); action representation is the axis B-spline Policy (B-spline Policy: Accelerating Manipulation Policies via B-spline Action Representations), A2A (A2A: Action-to-Action Flow Matching), and Mixture of Frames Policy (Mixture of Frames Policy: Multi-Frame Action Denoising for Bimanual Mobile Manipulation) have already staked out; the sim-to-real gap is what World Labs’ R2S2R (Building Worlds That Train Robots — Real-to-Sim-to-Real (R2S2R) as a Scalable Engine for Training and Evaluating Robot Policies) and LEGS (LEGS: Fine-Tuning Teleop-Free VLAs for Humanoid Loco-manipulation in an Embodied Gaussian Splatting World) target from opposite sides. The “world-action models without two GB200s per robot” observation aligns with the on-device WAM position TurboVLA (TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM) and Cosmos 3 Edge (Introducing Cosmos 3 Edge) already occupy. Not a new artifact — a useful outside-view sanity check that the internal taxonomy matches how the ecosystem is naming its problems.