Skip to content

Sunday Robotics ACT-2 Preview — unifying broad generalization with high reliability (Tony Zhao launch tweet)

Sunday Robotics co-founder Tony Zhao announces ACT-2 Preview, the successor to Sunday’s ACT-1 home-robot foundation model powering Memo. The pitch is a single sentence claim about the two axes the field usually treats as a trade-off: “the first robotics model to unify broad generalization with high reliability.” Two concrete numbers accompany the claim — a single fine-tuning example teaches Memo a new behavior that generalizes, and zero-shot deployment in real, unseen homes reaches 99% success rate. A 2:25 launch video is attached; no technical report, blog post, or paper is linked from the tweet itself.

  • ACT-2 Preview claims to unify broad generalization with high reliability in a single robotics foundation model [tweet].
  • A single fine-tuning example is claimed sufficient to teach the Memo robot a new behavior that then generalizes [tweet].
  • Zero-shot deployment in real unseen homes is claimed at 99% success rate [tweet].

Not disclosed in the tweet. Prior public context on Sunday: ACT-1 (November 2025) was trained on ~10M chore episodes from >500 U.S. homes collected via a 200200–400 Skill Capture Glove whose 3-finger geometry matches Memo’s grippers; a Skill Transform pipeline converts glove trajectories to robot actions with reported ~90% fidelity; ACT-1 is presented as trained on “zero robot data” and is described as the first end-to-end foundation model unifying long-horizon manipulation with map-conditioned navigation, demonstrated in six unseen Airbnb homes. Whether ACT-2 keeps that recipe unchanged and layers a fine-tuning mechanism on top, or restructures the pretraining/post-training stack, is not disclosed. The 99% number’s task, denominator, and evaluation protocol are not disclosed.

Two announced numbers with no methodology:

  • 99% zero-shot success rate in real unseen homes (task and denominator unspecified) [tweet].
  • Generalization from a single fine-tuning example (behavior and generalization axis unspecified) [tweet].

For anchoring, the strongest currently-filed comparable data point is π0.5’s ~40% real-robot success on the phone-packing / printer-refilling / laundry-loading / box-packing suite at <10 h/task of teleop demos (Xiaomi-Robotics-1 (XR-1) — Scaling VLA Foundation Models with 100K Hours of Embodiment-Free UMI Pre-training), which XR-1 lifts to ~75%. A 99% number in unseen homes, if it holds up under third-party evaluation, would sit well above every currently-filed generalist-manipulation result.

Sunday’s whole thesis — real-home glove-captured demonstrations rather than teleop or simulation — was a distinct-enough recipe in the VLA Models concept page that it deserves its own row; ACT-2’s headline number is the first time that recipe is claimed to jointly solve reliability and generalization, the two axes the concept page currently tracks as a spectrum across π*0.6 (action-pretraining), Spirit-v1.5 (clean-teleop-scaling), Embodied-R1.5 (unified-VLM-pointing), and XR-1 (embodiment-free UMI-scale). The “single fine-tuning example → generalizing behavior” hook is structurally reminiscent of RoboTTT — 8k-timestep robot policy via test-time training (Jim Fan / NVIDIA GEAR)‘s RoboTTT (one-shot in-context learning from a single human video via test-time training), but ACT-2 frames it as SFT-style fine-tuning rather than TTT — a different lever. Both claims (99% success in unseen homes, one-shot fine-tuning generalization) are strong enough to warrant skepticism until a report drops; treating the numbers as claim-level rather than measured is the right stance for now.