Skip to content

Egocentric-1M — largest egocentric video dataset (Build AI / Eddy Xu announcement)

Build AI (founded by 18-year-old Eddy Xu) announced Egocentric-1M on April 8, 2026, billed as the largest egocentric video dataset in the world and the third instalment in a rapid scaling progression: Egocentric-10K (Nov 2025, 10K hours, 2,153 workers) → Egocentric-100K (Dec 2025, 100K hours, 14,228 workers, 10.8B frames) → Egocentric-1M (~1M hours per third-party summaries). All entries are first-person factory-floor footage captured on Build AI’s custom head-mounted glasses across Southeast Asian manufacturing sites, released on Hugging Face under Apache 2.0. The tweet is positioned as “the next step in building the internet for physical AI” — egocentric video at internet scale as the pretraining substrate for embodied/VLA models, analogous to text scrape for LLMs.

  • The release is announced as the largest egocentric video dataset in the world, framed by Xu as “the internet for physical AI” [tweet body].
  • The prior Egocentric-10K instalment contained 10,000 hours from 2,153 factory workers and 1.08B frames at 16.4 TB, released on Hugging Face under Apache 2.0 (Nov 10, 2025) [quoted tweet].
  • The intermediate Egocentric-100K (Dec 2025) scaled to 100,000 hours from 14,228 workers wearing custom glasses ~7 h on average, ~11B frames, ~2M clips averaging 180 s each [third-party reports: Mike Kalil, Labellerr].
  • Data is collected exclusively in real factories across Southeast Asia (assembly, sorting, packaging, machining, quality inspection) using Build AI’s own manufactured smart-glasses hardware [Humanoids Daily, HyperAI summaries].
  • The dataset is positioned as direct competition to Ego4D (Meta, 3,670 h / 923 participants) and EPIC-KITCHENS (55 h / 32 participants), with an order-of-magnitude frame-count lead over both [OpenSource Press summary].

Not a research paper — there is no method beyond a data-collection pipeline. The mechanics, as reported across press coverage:

  • Build AI manufactures custom head-mounted recording glasses (after a failed US-manufacturing attempt, production was moved to Shenzhen).
  • Glasses are deployed in real factories across Southeast Asia; workers wear them during normal shifts (~7 h average wear time).
  • Footage captures hand movements, object interactions, task ordering, and how humans recover from errors — i.e. passive capture of skilled manual labor, not staged or scripted demonstrations.
  • Data is processed into clips (~180 s average for the 100K release), uploaded to Hugging Face under Apache 2.0, accessible via streaming (no full download required) and gated by an access form.

The thesis Xu states explicitly is that if LLMs learned from internet-scale text, physical AI will learn from internet-scale first-person video — and Build AI’s role is to produce that data at the relevant order of magnitude.

No model results — this is a dataset announcement. Adoption signal from third-party reports: the 100K release attracted ~2M pageviews and ~18K downloads within weeks of release. Company funding signal: 5MseedfromAbstractVentures,PearVC,HF0,withadditionalsupportfromZFellows;Xureportedlyturneddown>5M seed from Abstract Ventures, Pear VC, HF0, with additional support from ZFellows; Xu reportedly turned down >25M in equity offers. Scaling cadence claimed: 10K → 100K → 1M hours in five months (Nov 2025 → Apr 2026).

For Luma’s world-model and embodied-AI tracks, an Apache-2.0, internet-scale, first-person real-factory video corpus is the closest thing to a “Common Crawl for physical AI” that exists in the open today — and it’s larger than Ego4D by 2-3 orders of magnitude. Complements Unitree open-sources UnifoLM-WBT-Dataset — humanoid whole-body teleoperation dataset on the data-corpus side: UnifoLM-WBT is whole-body teleop trajectories on a humanoid platform (joint states + actions), while Egocentric-1M is human first-person video without robot actions — together they bracket the “demonstration data” question (humanoid-embodied with actions vs. human-embodied without actions but at much larger scale). Also relevant to Action100M: A Large-scale Video Action Dataset which scales VLM-labeled action data on internet videos; Egocentric-1M is the raw first-person counterpart, and the open question is whether VLM-derived dense captioning over Egocentric-1M would close the gap with action-labelled corpora. The rapid 10K → 100K → 1M scaling cadence is itself a signal worth tracking — if real, the 100× growth in five months matches no other open data release in 2026.