Skip to content

Figure launches Index — crowdsourced app-based real-world video dataset for humanoid robots

Figure comes out of stealth with Index, an app-based real-world data collection pipeline for humanoid-robot training. In four months the program has crossed 264,000 app downloads across 108 countries, ingested over 16M video uploads at >30 minutes of video per second, paid out 15Mtocontributors,andreached43,000+weeklyactiveusers.Adcockframesitasananswertothe"internetdoesntcontainrobotdata"problemrealworldactiondatahastobecollectedfromthephysicalworldandcommits15M** to contributors, and reached **43,000+ weekly active users**. Adcock frames it as an answer to the "internet doesn't contain robot data" problem — real-world action data has to be collected from the physical world — and commits **1B in spend over the next 12 months on data + compute, with a 100× scale-up target. Figure positions Index as the largest useful robot training dataset in the world and as the “groundwork for ordering robots as a service.”

  • Index is a Figure-exclusive data-collection app, built over the last four months and now out of stealth; the tweet presents raw usage numbers rather than a technical dataset schema [tweet, thread items 1–4].
  • 264,000 app downloads, 16M+ video uploads, >30 minutes of video ingested per second, 43,000+ weekly active users, $15M paid out to contributors, 108 countries covered, with a live-updating global map on the Figure website [thread item 3].
  • Positioning claim: Index is the largest useful robot training dataset in the world; a “detailed write-up” is referenced but not linked from the tweet thread (the linked follow-up post is not directly extractable at filing time) [thread item 4].
  • Forward commitment: Figure is on a path to 100× the current scale and will spend >$1B in the next 12 months on data and compute; Index is framed as groundwork for a “robots as a service” ordering flow [thread item 5].

The details of Index’s data pipeline are not in the tweet — only the top-level architecture is stated: contributors download an app, record real-world video (implicitly human-egocentric or human-perspective footage per Figure’s earlier public statements about Helix Lab and Helix training), upload it, and are paid. The referenced write-up is not fetchable from the tweet at filing time. What is publicly disclosed is (a) global geographic distribution as a coverage lever, (b) throughput-first framing (video-seconds ingested per real-time second is the headline metric), and (c) direct monetary incentives as the collection primitive.

There are no scientific results — this is a launch announcement with usage-metric receipts. The relevant numbers are the counts above (264k downloads, 16M videos, 30 min/s ingest rate, 15Mpaid,108countries,43kWAU)andtheforward15M paid, 108 countries, 43k WAU) and the forward 1B / 100× commitment.

Index is the most explicit industrial-scale bet on paid-crowdsourced human video as robotics pre-training substrate filed so far. It sharpens a growing pattern the wiki has been tracking: Dyna-2: A 1-Million-Hour Scaling Law for World-Action Models scales egocentric human video to 1M hours (~170 human-years) and reports the first cross-embodiment scaling law from human-only pre-training; HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining argues egocentric human data can outperform real-robot data at matched volume; Lightwheel AI open-sources EgoSuite-Open100K — 100K-hour egocentric human dataset with hand and body pose open-sources EgoSuite-Open100K (100k hours, ~10× larger than any prior egocentric open release) and RekaDaily-10k: Collecting 10,000+ Hours of Egocentric Household Manipulation Data collects 10k+ hours of first-person household manipulation. Figure’s Index is the first filed effort explicitly productizing the crowdsourced collection as an app with global reach and cash payouts — a distinct point in the design space vs research-lab collection (Reka), open community drops (Lightwheel), or lab-internal Helix Lab. It complements Going Beyond World Models & VLAs‘s “GEN-0 grows with physical interaction data” thesis by staking out the “collect the physical-interaction data first, everything else follows” position from the largest funded humanoid company.

The 30-min-of-video-per-second ingest rate implies ~1.8k video-hours per real-time hour — if sustained, that reaches Dyna-2’s 1M-hour scale in ~23 real-time days of collection at current tempo. Whether this scale delivers Helix (Figure’s VLA system) gains proportional to Dyna-2’s scaling-law extrapolation is the open empirical question, and it will be worth watching whether Figure follows the World Foundation Models “video is a first-class pretraining axis” position or the more classical VLA teleop+action-head recipe.