Skip to content

Lightwheel AI open-sources EgoSuite-Open100K — 100K-hour egocentric human dataset with hand and body pose

Lightwheel AI announced EgoSuite-Open100K on Aug 21, 2026 — 100,000 hours of fully-annotated egocentric human footage released on Hugging Face, positioned as “the largest fully annotated open egocentric human dataset.” The first 10,000 hours are live, with the remainder rolling out in stages. Coverage spans 15,000+ tasks across 15,000+ real scenes (factory floors, retail backrooms, and other work environments), with per-frame hand pose, body pose, subtask-level semantics, and wrist-camera coverage on a subset. Notably licensed for commercial training rather than research-only, and framed by Lightwheel as “the first public layer” of the data infrastructure they build for physical AI.

  • The release is billed as the largest fully-annotated open egocentric human dataset, with 100,000 hours across 15,000+ tasks and 15,000+ real scenes [tweet body].
  • The corpus is fully annotated with hand pose, body pose, and subtask-level semantics; a subset additionally has wrist-camera coverage [tweet body].
  • License permits commercial training, not just research use — a departure from many prior egocentric releases [tweet body].
  • The first 10,000 hours are available today via Hugging Face, with the remaining 90,000 hours to roll out in stages [tweet body].

Not a research paper — this is a dataset announcement. From the tweet:

  • Capture domain: real work environments across factory floors, retail backrooms, and analogous settings — first-person / egocentric video.
  • Annotations shipped alongside raw video: hand pose, body pose, subtask-level semantics, per-frame; wrist-camera coverage on part of the set (secondary viewpoint).
  • Distribution: Hugging Face, initial 10K-hour drop with staged expansion to 100K.
  • Framing: Lightwheel positions itself as data infrastructure for physical AI, calling Open100K “the first public layer” of that stack. The pitch — physical AI has its own scaling law and human data is the input — matches the framing used by other 2026 egocentric releases.

No model / no benchmark numbers — this is a raw dataset drop. The measurable claims are scale-side: 100K hours, 15K+ tasks, 15K+ scenes.

Slots into the crowded “at-scale egocentric human video corpus” landscape that has crystallized as the current front-runner substrate for physical-AI pretraining. Sits between Egocentric-1M — largest egocentric video dataset (Build AI / Eddy Xu announcement) (Build AI’s Egocentric-1M, ~1M hours factory-floor footage under Apache 2.0, described as “the internet for physical AI”) and RekaDaily-10k: Collecting 10,000+ Hours of Egocentric Household Manipulation Data (RekaDaily-10k, 10K h household egocentric under Apache 2.0) on the scale axis — 100K hours puts Lightwheel at roughly Egocentric-100K’s scale, but with fully-annotated hand/body pose and subtask semantics rather than raw video, addressing the “labels-on-real-video is the bottleneck, not raw collection” observation from HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining. The workplace-scene focus (factory floors, retail backrooms) is also a distinct axis from RekaDaily’s household coverage and closer to Build AI’s Southeast-Asia-factory sourcing — commercial-training license makes it especially relevant for teams that can’t ship models trained on research-only corpora. Complements From Foundation to Application: Improving VLA Models in Practice (LingBot-VLA 2.0) (LingBot-VLA 2.0’s 60K-hour mix of robot + egocentric human) and EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data (EgoScale’s 20,854h pseudo-labeled egocentric pretraining) on the supply side — those recipes have consumed similar-shape data at similar scale, but usually with pseudo-labels rather than fully-annotated ground truth.