Skip to content

Sunday Robotics ran the Chinchilla scaling law with their VLA model (Tony Zhao follow-up)

In-thread reply to Pratyush Ranjan Tiwari’s comment that a robotics scaling-law result would be “as significant as the chinchilla-scaling-law moment,” Sunday Robotics co-founder Tony Zhao (Sunday Robotics ACT-2 Preview — unifying broad generalization with high reliability (Tony Zhao launch tweet)) writes that they “actually ran the exact chinchilla-scaling-law with our model” and defers a public writeup to @nadeesha99. No numbers, plot, or methodology are disclosed — just an assertion that Chinchilla-style compute-optimal scaling was performed on Sunday’s VLA. Filed as a marker for the claim while the underlying result remains unpublished.

  • Sunday Robotics ran the “exact chinchilla-scaling-law” on their VLA model [tweet].
  • A public writeup / talk is deferred to @nadeesha99 at Sunday, with no committed date [tweet].

Not disclosed. “Chinchilla scaling law” in the LLM literature refers to Hoffmann et al.’s isoFLOP procedure: sweep model size NN against training tokens DD at fixed compute budgets, fit a joint parametric loss L(N,D)L(N, D), and derive the compute-optimal (N,D)(N^*, D^*) pair — with the headline finding that compute-optimal DND^* \propto N (i.e. constant tokens-per-parameter). What the analog looks like for a VLA is exactly the open question the tweet skips: the “data” axis could be teleop-episode-hours, action-labeled human-video-hours, glove-capture-episodes, or a hybrid; the loss could be a validation action-prediction loss (as EgoScale and XR-1 both use for their scaling curves) or a real-robot success rate proxy. None of that is specified in the tweet.

None disclosed.

If the claim holds up in a full report, this would be the first filed instance of Chinchilla-style isoFLOP scaling being run on a production home-robot VLA — sharper than the log-linear loss-vs-hours fits EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data and Xiaomi-Robotics-1 (XR-1) — Scaling VLA Foundation Models with 100K Hours of Embodiment-Free UMI Pre-training have already put on the Hyperparameter scaling laws page, because a joint (N,D)(N, D) fit gives a compute-optimal recipe rather than just a data-axis power law. It would also directly test whether AR LLMs’ DND^* \propto N result transfers to VLA training — The Design Space of Tri-Modal Masked Diffusion Models already showed the equivalent for tri-modal masked diffusion gives DN0.476D^* \propto N^{0.476} (larger models more data-hungry per parameter), so it is genuinely open which regime a VLA falls in. Sits in the same reliability-vs-generalization thesis as Sunday Robotics ACT-2 Preview — unifying broad generalization with high reliability (Tony Zhao launch tweet) (the ACT-2 Preview launch tweet this reply is threaded to) — Sunday’s implicit argument is that their real-home glove-capture data substrate scales predictably enough to admit a Chinchilla-style fit, distinguishing their recipe from teleop-only competitors. Treat the claim as unverified until the writeup lands.