TrueLabel — Physical AI Dataset Directory: Robotics, Humanoid & Egocentric Data
TrueLabel’s Physical AI Dataset Directory is a browsable catalog claiming 658 indexed dataset pages across robotics, humanoid, egocentric, teleoperation, manipulation, and simulation data. Each row carries provider, year, license, modality coverage, robot platform, popularity signal (download count / likes / size bucket), and commercial-use guidance, and links out to the source page (typically Hugging Face). The visible top-50 listing skews heavily toward 2026-04 / 2026-05 / 2026-06 Hugging Face LeRobot-format releases (Bridge / DROID re-uploads, AgiBot World 2026, RoboCasa365, DROID 1.0.1 mirrors, NVIDIA GR00T X-Embodiment Sim, OmniWorld, NVIDIA NuRec, BONES-SEED, MuscleMimic, …) alongside the canonical pre-2026 foundations (Open X-Embodiment, BridgeData V2, RT-1, ALOHA, RoboMimic, Meta-World, RLBench, CALVIN, ManiSkill, Ego4D, EPIC-KITCHENS, HOI4D, DexYCB, ScanNet, Habitat, AI2-THOR, BEHAVIOR, Waymo, Kinetics, RoboSuite, RH20T, LIBERO, RoboTurk, UMI, FurnitureBench, TACO Play).
Key claims
Section titled “Key claims”- Indexes 658 robotics / embodied-AI dataset pages with per-entry provider, year, license, modality, robot platform, and procurement notes — a single browsable surface across Hugging Face, NVIDIA, academic releases, and benchmark-suite repositories [landing page header + entries].
- Tags each entry with commercial-use guidance directly in the row: explicit “Commercial use unclear”, “Commercial use restricted”, “Source appears permissive; verify data terms” labels appear repeatedly, so license risk is visible at browse time rather than at download time [landing page entries].
- Surfaces popularity proxies per entry — Hugging Face download counts (e.g. AgiBotWorld2026 47,447; OmniWorld 200,677; 10Kh-RealOmin-OpenData 200,660; xperience-10m 188,007) and per-entry “likes” — letting a reader sort by adoption signal [landing page entries].
- Distinguishes simulation from real-world capture inline: entries flag “simulated manipulation benchmark” (RLBench, Meta-World, CALVIN, ManiSkill, RoboSuite, RoboCasa365, vlabench), simulated trajectories (NVIDIA GR00T X-Embodiment Sim, “9,000 simulated bimanual manipulation trajectories”), and synthetic vs. teleop demonstrations (RoboCasa “600+ hours of human demonstration data, and 1,600+ hours of synthetic demonstrations”) [landing page entries].
- “Best for” hint per entry maps the dataset to a downstream training objective (“robot foundation model pretraining”, “behavior cloning”, “VLA benchmark evaluation”, “egocentric perception”, “language-conditioned policy evaluation”, “dexterous grasping”, “robotics dataset distribution”) [landing page entries].
Method
Section titled “Method”The site is a curated index, not a host: each row reports metadata pulled from the source provider (license, download count, robot platform, modality) and links out to the Hugging Face / GitHub / project page that actually holds the data. The 658-page claim is taken at face value from the landing-page header; the browseable surface returns 50 results at a time. The entries themselves are short summaries — usually one sentence of dataset description, plus a quoted line from the source describing scale (“219.6 hours and 5,100 successful furniture assembly demonstrations”, “30,050 trajectories, including 9,500 collected through teleoperation”, “1M+ trajectories and 2,976.4 hours across 217 tasks”) — followed by the commercial-use flag and “Best for” hint. Curator authorship is undisclosed on the landing page.
Results
Section titled “Results”Not a research artifact — no benchmark numbers to report. The catalog itself is the artifact. As of 2026-06-30, the visible top-of-list bias is heavily toward Hugging Face LeRobot v3.0 community releases (the “This dataset was created using LeRobot. Dataset Structure meta/info.json” boilerplate appears on dozens of entries), with a long tail of canonical benchmark suites (Open X-Embodiment, RT-1, RLBench, Meta-World, …). Modality tagging is implicit (license + robot platform + dataset description) rather than as a structured tag set like datasets.bot uses, but the commercial-use column is more aggressively labeled — “Commercial use restricted” appears on the Ego4D / EPIC-KITCHENS / Kinetics / Waymo / Something-Something V2 / ScanNet entries, which is exactly the set that breaks the most commercial pipelines downstream.
Why it’s interesting
Section titled “Why it’s interesting”TrueLabel is the third-party-catalog sibling of datasets.bot — Curated catalog of robotics and embodied-AI training datasets — same artifact pattern (browsable index of robotics/embodied datasets with license + modality + popularity metadata), different curator and roughly an order of magnitude more entries (658 indexed pages vs datasets.bot’s landing surface). It contrasts with Sophon — AI Research Catalog of Evals, Tools, and Labs (evals/tools/labs catalog) and awesome-diffusion-papers.vercel.app — time-sorted, searchable browser for awesome-diffusion-categorized (diffusion-paper catalog): same third-party-catalog pattern, but specifically for the embodied-data substrate. Useful as a scoping dashboard when the recurring “should we train on dataset X or Y?” question hits the team — particularly for the egocentric-vs-teleop pretraining debate that Human-to-Robot Retargeting tracks (HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining, EgoInfinity: A Web-Scale 4D Hand-Object Interaction Data Engine for Any-View Robot Retargeting and Video-to-Action Robot Learning, What Matters When Cotraining Robot Manipulation Policies on Everyday Human Videos?) — TrueLabel’s combined view of Ego4D / EPIC-KITCHENS / HOI4D / DexYCB / xperience-10m / 10Kh-RealOmin-OpenData / UniHand_Preview alongside the AgiBot World, DROID, and LeRobot teleop releases is the closest thing to a side-by-side license-and-scale comparison currently filed.
See also
Section titled “See also”- datasets.bot — Curated catalog of robotics and embodied-AI training datasets — direct sibling: same third-party-dataset-catalog pattern, smaller visible surface but cleaner explicit modality vocabulary.
- Sophon — AI Research Catalog of Evals, Tools, and Labs — same third-party-catalog pattern for evals/tools/labs.
- awesome-diffusion-papers.vercel.app — time-sorted, searchable browser for awesome-diffusion-categorized — same pattern for diffusion papers.
- Survey notes on World Action Models / VLA for robotics (index) — narrative survey of the WAM/VLA corpus; TrueLabel is the dataset-side complement.
- VLA Models — the dominant training target for the catalog’s entries.
- Human-to-Robot Retargeting — the bridging-data debate this catalog can be used to scope across egocentric, teleop, and simulated substrates.
- Synthetic Training Data — TrueLabel indexes synthetic releases (RoboCasa synthetic demos, NVIDIA GR00T X-Embodiment Sim, vlabench_primitive_ft_lerobot_video) alongside real-world captures.
- World Foundation Models — the catalog is useful for sourcing the perception substrate WFMs are trained on.