Skip to content

Introducing Index: Building The World's Largest and Most Diverse Physical Dataset

Figure comes out of stealth with Index, an app-based crowdsourced pipeline for collecting real-world human video for humanoid-robot training. Beyond the top-line numbers previously teased on Twitter (264k downloads, 108 countries, 44k WAU, 16M videos, 30 min/s ingest, 15Mpaidout),theblogdisclosestheactualfivestagepipelinefiltering,fraudreview,deduplication,rebalancing,annotationanddiversitymetrics:373uniquetasks,1,146uniquemanipulatedobjects,and116uniqueenvironmentsper1,000hourscollected.Figureframesthisasa15M paid out), the blog discloses the actual **five-stage pipeline** — filtering, fraud review, deduplication, rebalancing, annotation — and diversity metrics: **373 unique tasks, 1,146 unique manipulated objects, and 116 unique environments per 1,000 hours** collected. Figure frames this as a 1B/12-month bet that “generalization is a data problem” for Helix and positions Index as groundwork for a future “robots as a service” ordering flow.

  • Index is Figure’s response after “vendors couldn’t hit the throughput, diversity, or quality bar Helix requires” — an internally-built, Figure-exclusive collection app now on Google Play and the App Store [§Index].
  • Ingest scale: 30 minutes of video per second, equivalent to 4.9 years of human work uploaded per day; 264k downloads across 108 countries, 44k+ WAU, 16M videos, $15M paid out to Creators [§Index].
  • Per-1,000-hour diversity: 373 unique tasks, 1,146 unique manipulated objects, 116 unique environments — diversity comes from Creator heterogeneity (“every new Creator brings an unseen environment, unfamiliar objects, and their own idiosyncratic way of completing a task”) [§Index].
  • Task coverage spans cooking, cleaning, laundry, household chores, and business settings (logistics centers, restaurants, factories, offices), with named long-tail examples like changing oil, cleaning kitty litter, and busing tables [§Index].
  • The pipeline has five stages: (1) automated technical/visual/semantic quality filters, (2) human fraud audits at the user level, (3) embedding-based near-duplicate removal, (4) task-quota + embedding-cluster rebalancing, and (5) hierarchical text captioning of each episode [§Using Human Data for Helix Training].
  • Forward commitment: 100× current scale and >$1B spend on data + compute over the next 12 months; Index is framed as groundwork for a “robots as a service” ordering flow [§Next Steps].
  • “Generalization is a data problem” — Figure claims internal generalization results validate the thesis, with details promised in a follow-up post [§Generalization is a Data Problem].

The disclosed data pipeline is five sequential stages:

  1. Filtering — automated screens for technical quality (encoding, resolution, motion), visual quality, and semantic quality (does the clip match a real task).
  2. Fraud review — human analysts audit at the user level to catch coordinated evasion of automated filters (implying continuous adversarial-quality-control loops).
  3. Deduplication — each video segment is embedded, and segments above a similarity threshold to already-accepted data are discarded, protecting diversity as the corpus grows.
  4. Rebalancing — surviving segments are re-weighted using two signals: (a) how well a submission matches the Creator’s selected task label (task-quota alignment) and (b) embedding-based clusters that capture variation beyond task labels.
  5. Annotation — hierarchical text captions are generated for every episode.

Steps 3 and 4 together constitute the diversity engine: deduplication keeps the raw stream from collapsing to popular tasks/environments, and cluster-level rebalancing over embeddings actively rewards long-tail contributions. Fraud review at the user level (not the clip level) is the interesting operational detail — it suggests Figure is optimizing for Creator-population honesty rather than per-video judgments, which likely matters as the payout budget scales into the hundreds of millions.

The data-collection primitive is (a) a mobile app (Google Play + App Store), (b) an optional recording device shipped to signed-up Creators, and (c) a marketplace where Creators can either record their own household/workplace activity or be booked to visit paying customers’ homes and businesses. Paid-per-minute contribution model, with 15MpaidoutoverfourmonthsimplyinganaveragepayoutratethatFigurehasnotpubliclydisclosedbutthattriviallydividestoroughly15M paid out over four months implying an average payout rate that Figure has not publicly disclosed but that trivially divides to roughly 3.75M/month at current tempo.

  • Downloads: 264,000 across 108 countries.
  • Weekly active users: 44,000+.
  • Total videos uploaded: 16M+.
  • Ingest rate: 30 minutes of video per second (≈1,800 video-hours per real-time hour, or 4.9 human-years per calendar day).
  • Payouts: $15M to Creators cumulatively.
  • Diversity per 1,000 hours: 373 unique tasks, 1,146 unique manipulated objects, 116 unique environments.
  • Forward commitment: 100× scale-up target and >$1B on data + compute in the next 12 months.

No policy-training results, benchmark numbers, or scaling curves are shared in this post — Figure explicitly promises those “in detail soon.”

Index is the most explicit industrial-scale bet on paid-crowdsourced human video as robotics pre-training substrate filed on the wiki, and this blog post now discloses the pipeline mechanics that the launch tweet (Figure launches Index — crowdsourced app-based real-world video dataset for humanoid robots) only hinted at. The five-stage pipeline (esp. embedding-based dedup + cluster-level rebalancing) is a concrete answer to the objection that raw crowdsourced video will collapse toward frequent, easy tasks — Figure is treating diversity as a first-class quantity to be actively optimized, not just a byproduct of Creator count.

The disclosed per-1,000-hour diversity metrics (373 tasks / 1,146 objects / 116 environments) are the first filed quantitative diversity claims for a robot-scale human video corpus. For comparison, RekaDaily-10k: Collecting 10,000+ Hours of Egocentric Household Manipulation Data collects 10k+ hours of first-person household manipulation via a research-lab process, and Lightwheel AI open-sources EgoSuite-Open100K — 100K-hour egocentric human dataset with hand and body pose open-sources EgoSuite-Open100K (100k hours) — but neither reports task/object/environment cardinality on a normalized-per-hour basis, which makes Figure’s numbers the first head-to-head-comparable diversity benchmark for future entrants.

The pipeline design also contrasts sharply with Dyna-2: A 1-Million-Hour Scaling Law for World-Action Models (Dyna’s 1M-hour scaling law): Dyna scales via quantity on YouTube-scraped web video and derives a cross-embodiment scaling law from it, while Figure is spending real money to control diversity and quality on an owned collection channel. The interesting empirical question in six months is whether Dyna-2’s scaling extrapolation holds for controlled-diverse app-collected data at the same total hours, or whether Figure’s per-hour diversity multiplier changes the exponent.

Finally, the “generalization is a data problem” framing (§Generalization is a Data Problem) puts Figure on record with the Going Beyond World Models & VLAs / GEN-0 school — physical-interaction data scale is the primary lever, not architectural novelty. This is now the second billion-dollar-scale filed commitment to that thesis after Generalist AI’s public roadmap, and the first with app-crowdsourced collection as the mechanism.