Skip to content

York Yang (Dyna Robotics) — ROI and scalability are inseparable; every deployment must make the next one better

York Yang (co-founder Dyna Robotics, ex Caper AI / Instacart) posts a self-contained positioning thread arguing that the load-bearing metric for robotics is not demo success but scalable ROI: every deployment must make the next one better, or the economics never compound. He frames the path as a three-step ladder — Demo (“it works”) → Pilot (“it works here”) → Scaled deployment (“we want more”) — and insists that research and deployment must stay tightly coupled, because deployment is what tells the research side what actually matters and where things break. No numbers, no product announcement, no linked paper — a company-thesis post about which R&D axis is load-bearing for commercial robotics.

  • The bar for robotics has moved past “impressive demo” — the real test is whether customers get enough value that they want to deploy more [thread ¶1].
  • ROI and scalability are inseparable: a single successful deployment can prove customer ROI, but not a scalable product or business; if every new customer or workflow requires rebuilding the solution, the economics never scale [thread ¶3].
  • Scalable ROI is only proven when the same underlying technology creates value across customers, workflows, and industries while the effort/cost/time per new deployment keeps declining [thread ¶3].
  • Three-step deployment ladder: Demo (“it works” — technical possibility) → Pilot (“it works here” — value in a real environment) → Scaled deployment (“we want more” — customers expand because ROI works and the vendor can keep delivering without rebuilding) [thread ¶4].
  • Every deployment must make the next one better — a problem solved in the field should leave something reusable behind: a better model, better tooling, better infrastructure, or a more general capability [thread ¶5].
  • Research and deployment must stay tightly connected: research expands what robots can do; deployment reveals what actually matters, where things break, and what must be solved fundamentally rather than patched case by case [thread ¶6].
  • Explicit statement of Dyna’s product/research priority ordering: robots must create real value, that value must repeat, and it must scale [thread ¶8].

Not applicable — this is a positioning thread, no methodology or artifact.

Not applicable — no numbers, no benchmarks, no released code or model. The thread teases that Dyna is “sharing more of what we’ve learned across a broad range of industry partners — the successes, failures, operational challenges, and hard-earned lessons” but the tweet itself is announcement + frame only.

This is the fourth positioning artifact filed in the last four weeks where a robotics operator publicly names the deployment economics loop — not the model layer — as the load-bearing R&D axis. Sits alongside Robo Robotics launch — Robo-T bimanual humanoid, sub-$10/hour with Roboport teleoperation platform (Kyle Noble) (Robo Robotics’ teleoperator→supervisor→rare-intervener transition + Roboport DAgger substrate), MicroFactory — 99.9% reliability via $5 human-in-the-loop DAgger retraining on Jetson (Ilir Aliu × Igor Kulakov podcast) (MicroFactory’s 99.9% reliability via $5 human-in-the-loop DAgger retraining), and Enact launch — post-training infrastructure that generates targeted recovery data for robotics VLAs (Enact’s post-training infrastructure that generates targeted recovery data for VLAs). Yang’s framing is the most abstract of the four — he does not commit to a specific technical lever (teleop economics, intervention rate, DAgger loop) — but his three-step ladder and “every deployment must make the next one better” reusability constraint is the underlying business-model claim all four converge on.

Dyna Robotics is one of the most-filed labs on the wiki (see Dyna-2: A 1-Million-Hour Scaling Law for World-Action Models Dyna-2’s 1M-hour WAM scaling law, Training Dyna-2 at million-hour scale, repeatably its infra companion, plus prior Dyna Robotics tweets), so this thread is best read as the company’s articulation of why it made the technical bets it did — pretraining on ordered-of-magnitude more human video, running WAM-vs-VLA head-to-heads under matched conditions, publishing customer-site deployment pass rates (46% → 87%) rather than only in-distribution completion rates. The 46/87 gap on the deployment side vs matched ~100% completion in the pretraining paper is a concrete instance of exactly what Yang is arguing here: completion-only metrics don’t measure the property that decides whether the customer wants more.