RoboArena rolls back evaluations after benchmark hacking observed since April (Pranav Atreya announcement)
RoboArena’s lead author Pranav Atreya announces that the team observed evidence of benchmark hacking on the RoboArena leaderboard dating back to April, has put new safeguards in place, and has rolled back affected evaluations to preserve benchmark integrity — pointing to robo-arena.github.io for details on the changes. The slack pointer from @amit adds the key piece of context: NVIDIA’s Cosmos was among the entries removed as part of the rollback. The announcement lands ~12 days after Cosmos 3 was launched at Computex Taipei (June 1) claiming top RoboArena placement, and ~10 days after Spirit v1.6 was widely reported as displacing Cosmos3-Nano-Policy at the top of that same leaderboard — making this the first public acknowledgement from RoboArena’s maintainers that the post-Cosmos-3 leaderboard movement was non-trustworthy.
Key claims
Section titled “Key claims”- Evidence of benchmark hacking on RoboArena was observed beginning April 2026 [tweet body].
- The RoboArena team has implemented changes to prevent future hacking and rolled back affected evaluations [tweet body].
- A details page at
robo-arena.github.iois referenced as the canonical writeup of the changes [tweet embedded URL].
Method
Section titled “Method”This is a 1-sentence public statement on X from Pranav Atreya (first author of the RoboArena paper RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies) announcing an operational decision. No technical mechanism for either the detected hacking or the implemented mitigation is given in the tweet itself — the linked landing page is referenced as the substantive writeup but returned empty content at fetch time, so the specifics (which policies were rolled back, what the new evaluator-side or submitter-side safeguards are, how submissions back to April are being audited) are not retrievable from the tweet alone. The Slack-pointer note from @amit is the only public signal in the wiki that Cosmos was among the removed entries.
Results
Section titled “Results”No quantitative results. The single operational outcome is the rollback itself: leaderboard entries from “since April” are no longer treated as part of the public ranking, and at least one specific entry (Cosmos, per @amit) is removed.
Why it’s interesting
Section titled “Why it’s interesting”This is a live integrity event on the wiki’s most heavily cited robot-policy benchmark, and it directly punctures three previously filed claims. First, NVIDIA Launches Cosmos 3, the Open Frontier Foundation Model for Physical AI and How Cosmos 3 Helps Physical AI Think Before It Acts both lean on the Cosmos 3 Nano post-trained policy “leading RoboArena” as the headline policy-benchmark datapoint — that claim is now retracted by the benchmark maintainer; the wiki’s Cosmos 3 pages should be read with that caveat. Second, Spirit AI beats Nvidia on RoboArena robotics benchmark reports Spirit v1.6 unseating Cosmos3-Nano-Policy at 1,924 vs 1,881 within 48 hours of Cosmos 3’s launch; the rollback retroactively casts doubt on the entire post-Cosmos-3 leaderboard ordering, not just Cosmos’s slot. Third, RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies‘s design pitch was that double-blind crowd-sourced pairwise comparison is the trust mechanism that fixes-task-set benchmarks lack; an actually-detected hacking attack on that design is the first filed counterexample, and a serious one — it means the §3 four-property argument (robustness against single-actor manipulation in particular) survived a real adversarial test only by the maintainer team noticing out-of-band and intervening manually, not by the protocol catching it automatically.
For RL Environment Platforms this is the first filed integrity-attack datapoint on any env platform in the cluster: SETA, Toolathlon-GYM, OpenReward, RoboCasa365, Genesis World have all been filed as architectural artifacts, with no filed evidence of in-the-wild adversarial use yet. For VLM-as-Evaluator it raises the question of whether the §3.3 GPT-4.5 + o3 qualitative-analysis pipeline contributed to detection or missed it. For Open foundation-model releases and World Foundation Models this is direct evidence that “open frontier WFM” leaderboard claims need a separate trust check — the rollback removes a specific high-profile entry without (in the tweet) naming it.
See also
Section titled “See also”- RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies — the benchmark whose integrity event this announces; Atreya is first author
- NVIDIA Launches Cosmos 3, the Open Frontier Foundation Model for Physical AI — Cosmos 3 launch post claiming top RoboArena placement, now retracted
- How Cosmos 3 Helps Physical AI Think Before It Acts — companion NVIDIA blog reiterating the Cosmos 3 Nano RoboArena lead claim
- Cosmos 3: Omnimodal World Models for Physical AI — Cosmos 3 technical report
- Spirit AI beats Nvidia on RoboArena robotics benchmark — news piece reporting Spirit v1.6 topping Cosmos3-Nano on the same leaderboard 48h after launch
- RL Environment Platforms — first filed integrity-attack datapoint on any platform in the cluster
- VLM-as-Evaluator — RoboArena’s analysis layer; relevant to detection question