Sunday Robotics data engine — from one Memory Developer off Craigslist to 1,000+ (Perry Jia thread)
Perry Jia — who leads Data Operations at Sunday Robotics after ~6 years running Tesla’s data-engine programs on Autopilot and Optimus — opens a thread claiming Sunday’s ACT-2 data engine is now backed by 1,000+ paid “Memory Developers” recruited over two years, starting from a single hire off Craigslist onboarded in a public library. Sunday’s public-facing framing of the program is that Memory Developers use provided hardware (the Skill Capture Glove tied to Memo’s 3-finger gripper geometry) to record real household demonstrations in their own homes at 60/hour of approved data. The thread is a personnel/operations claim about the labor substrate under Sunday’s real-home glove-capture recipe, not a technical writeup. Filed as a marker for the scale claim — the follow-up posts (2/N…) that presumably carry method and numbers are not resolvable from the retrieved tweet content.
Key claims
Section titled “Key claims”- Sunday’s Memory Developer program has grown from 1 hire (recruited off Craigslist, onboarded in a public library) to 1,000+ contributors over roughly two years [tweet 1/N].
- The Memory Developer workforce is described as the data engine behind ACT-2 — the labor substrate underneath the model, not a separate line of business [tweet 1/N].
Method
Section titled “Method”Not disclosed in the retrieved opening post; a 2:25 video is attached and the thread is threaded (1/N…) but the retrieved fetch does not surface the follow-up posts. Sunday’s publicly documented Memory Developer job posting describes the operational model: remote, part-time role using provided hardware (Skill Capture Glove) to record household-task demonstrations, screening period paid at 60/hour post-screening for complex tasks, “pay-per-task” rather than hourly. The ACT-2 Preview blog (ACT-2 Preview: Generalizing Reliability) describes the data engine as one of four end-to-end components (“the robot, the model, the fleet, and the data engine”) but does not disclose contributor count or per-hour economics.
Results
Section titled “Results”None disclosed in the opening post. External reporting places ACT-1’s training set at ~10M household episodes from >500 U.S. homes; ACT-2 Preview’s headline number is 99% zero-shot success in unseen homes (Sunday Robotics ACT-2 Preview — unifying broad generalization with high reliability (Tony Zhao launch tweet)). This tweet’s net-new datapoint is the labor-scale figure — 1,000+ Memory Developers behind that data — which has not been publicly stated at that precision before.
Why it’s interesting
Section titled “Why it’s interesting”The VLA Models concept page tracks multiple structural levers for generalist manipulation — action-pretraining (π*0.6: a VLA That Learns From Experience (RECAP)), clean teleop (Spirit-v1.5: Clean Data Is the Enemy of Great Robot Foundation Models), unified-VLM pointing (Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models), frozen-WFM + action expert (μ₀: A Scalable 3D Interaction-Trace World Model), scale-of-embodiment-free-UMI-data (Xiaomi-Robotics-1 (XR-1) — Scaling VLA Foundation Models with 100K Hours of Embodiment-Free UMI Pre-training) — but the operational lever underneath Sunday’s real-home glove-capture recipe has been visible only in aggregate (“~10M episodes from >500 homes”). This tweet sharpens the picture by putting a labor number on it: the recipe requires standing up a 1,000+-contributor paid workforce, roughly analogous to Tesla’s Autopilot data-labeling ops (which Perry Jia previously ran). That’s a distinct kind of “data lever” from mimic’s wearable-hardware pyramid (Solving Dexterity: A Full-Stack Approach (mimic hand M1 + wearable U1)) or XR-1’s 100K-hour UMI collection (Xiaomi-Robotics-1 (XR-1) — Scaling VLA Foundation Models with 100K Hours of Embodiment-Free UMI Pre-training) — Sunday is betting that a distributed paid workforce in real homes is what the recipe needs, and that this labor substrate is itself the moat, complementary to the Chinchilla-style scaling-law claim from the same thread family (Sunday Robotics ran the Chinchilla scaling law with their VLA model (Tony Zhao follow-up)). Reads best alongside Low-data inductive-bias lessons don't translate to the high-data regime — embracing chaos produced emergent policy capabilities (Nishant Desai / Sunday Robotics)‘s “embracing chaos” post from the same launch cycle — chaos in the training distribution has to come from somewhere, and this tweet names the source.
See also
Section titled “See also”- Sunday Robotics ACT-2 Preview — unifying broad generalization with high reliability (Tony Zhao launch tweet) — ACT-2 Preview launch tweet; this thread is Sunday’s operational-side companion, naming the labor engine behind the model announced there.
- Sunday Robotics ran the Chinchilla scaling law with their VLA model (Tony Zhao follow-up) — Chinchilla-scaling-law follow-up in the same launch cycle; the model-side scaling claim that this operational-side thread is the labor complement to.
- Low-data inductive-bias lessons don't translate to the high-data regime — embracing chaos produced emergent policy capabilities (Nishant Desai / Sunday Robotics) — “embracing chaos” post from Sunday’s launch week; chaos in the distribution requires the diverse real-home data-collection workforce this thread describes.
- ACT-2 Preview: Generalizing Reliability — Sunday’s official ACT-2 Preview writeup; frames “the data engine” as one of four end-to-end components but doesn’t publish workforce numbers.
- VLA Models — the recipe-lever concept page; Sunday’s real-home-glove-capture position gets an operational label here (distributed paid workforce) rather than only a data-source label.
- Synthetic Training Data — adjacent concept: Sunday’s data is real rather than synthetic, but the “paid contributors record real demonstrations at scale” workflow is a specific alternative to the synthetic-data recipes the concept catalogs.