Sunday Robotics ran the Chinchilla scaling law with their VLA model (Tony Zhao follow-up)
In-thread reply to Pratyush Ranjan Tiwari’s comment that a robotics scaling-law result would be “as significant as the chinchilla-scaling-law moment,” Sunday Robotics co-founder Tony Zhao (Sunday Robotics ACT-2 Preview — unifying broad generalization with high reliability (Tony Zhao launch tweet)) writes that they “actually ran the exact chinchilla-scaling-law with our model” and defers a public writeup to @nadeesha99. No numbers, plot, or methodology are disclosed — just an assertion that Chinchilla-style compute-optimal scaling was performed on Sunday’s VLA. Filed as a marker for the claim while the underlying result remains unpublished.
Key claims
Section titled “Key claims”- Sunday Robotics ran the “exact chinchilla-scaling-law” on their VLA model [tweet].
- A public writeup / talk is deferred to @nadeesha99 at Sunday, with no committed date [tweet].
Method
Section titled “Method”Not disclosed. “Chinchilla scaling law” in the LLM literature refers to Hoffmann et al.’s isoFLOP procedure: sweep model size against training tokens at fixed compute budgets, fit a joint parametric loss , and derive the compute-optimal pair — with the headline finding that compute-optimal (i.e. constant tokens-per-parameter). What the analog looks like for a VLA is exactly the open question the tweet skips: the “data” axis could be teleop-episode-hours, action-labeled human-video-hours, glove-capture-episodes, or a hybrid; the loss could be a validation action-prediction loss (as EgoScale and XR-1 both use for their scaling curves) or a real-robot success rate proxy. None of that is specified in the tweet.
Results
Section titled “Results”None disclosed.
Why it’s interesting
Section titled “Why it’s interesting”If the claim holds up in a full report, this would be the first filed instance of Chinchilla-style isoFLOP scaling being run on a production home-robot VLA — sharper than the log-linear loss-vs-hours fits EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data and Xiaomi-Robotics-1 (XR-1) — Scaling VLA Foundation Models with 100K Hours of Embodiment-Free UMI Pre-training have already put on the Hyperparameter scaling laws page, because a joint fit gives a compute-optimal recipe rather than just a data-axis power law. It would also directly test whether AR LLMs’ result transfers to VLA training — The Design Space of Tri-Modal Masked Diffusion Models already showed the equivalent for tri-modal masked diffusion gives (larger models more data-hungry per parameter), so it is genuinely open which regime a VLA falls in. Sits in the same reliability-vs-generalization thesis as Sunday Robotics ACT-2 Preview — unifying broad generalization with high reliability (Tony Zhao launch tweet) (the ACT-2 Preview launch tweet this reply is threaded to) — Sunday’s implicit argument is that their real-home glove-capture data substrate scales predictably enough to admit a Chinchilla-style fit, distinguishing their recipe from teleop-only competitors. Treat the claim as unverified until the writeup lands.
See also
Section titled “See also”- Sunday Robotics ACT-2 Preview — unifying broad generalization with high reliability (Tony Zhao launch tweet) — parent launch tweet in the same thread; ACT-2 Preview’s 99% real-home success + one-shot-example generalization claims are the capability side of the same recipe this follow-up puts scaling-law evidence behind.
- Hyperparameter scaling laws — where a full Sunday Chinchilla-style writeup would land; would be the first VLA joint isoFLOP fit if delivered.
- VLA Models — Sunday’s real-home glove-capture recipe sits as a distinct row in the recipe-lever debate; a scaling-law fit would sharpen its “data-substrate-as-lever” position vs Spirit-v1.5: Clean Data Is the Enemy of Great Robot Foundation Models (clean teleop) and Xiaomi-Robotics-1 (XR-1) — Scaling VLA Foundation Models with 100K Hours of Embodiment-Free UMI Pre-training (100K UMI hours).
- EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data — closest currently-filed VLA data-axis scaling result (log-linear loss vs hours, R²=0.9983), against which a Sunday Chinchilla fit would be the natural next step up.