EgoSuite-Open100K: 100,000 hours of egocentric human data for Physical AI
Lightwheel has released EgoSuite-Open100K, an open egocentric human activity dataset built for Physical AI, in partnership with Hugging Face. The full dataset will total 100,000 hours of first-person human activity across 15,000+ tasks and 15,000+ real-world collection scenes. The first 10,000 hours are available now.
Not sure where to start? EgoDemo is a 50-hour sample covering every annotated subset plus two raw-video variants — a fast way to get a feel for the data before pulling the full release.
At a glance
| Total dataset | 100,000 hours |
| Available now | 10,000 hours |
| Tasks | 15,000+ |
| Collection scenes | 15,000+ |
| Environmental categories | 7 |
| Scene types | 128 |
| Task categories | 18 |
| Annotations | Hand pose, body pose, plus event-level semantic annotation on select subsets |
| Camera setup | Egocentric head-mounted; wrist camera in EgoPro |
| Formats | LeRobot v3, MCAP |
| Usage | Academic research and commercial training |
Why egocentric human data
Robots don't just need to see the world, they need examples of how people act in it: reaching, grasping, sequencing steps, recovering from a fumble, finishing a task start to finish. That's the kind of supervision egocentric human video can provide at a scale robot-only data collection struggles to match. Robot-specific data can then build on top of it rather than carry the whole load alone.
There's growing evidence this scales. NVIDIA's EgoScale work found a log-linear scaling relationship across 20,854 hours of egocentric human video, and separate scaling experiments from Dyna Robotics point the same direction: more, and more diverse, human experience keeps improving downstream performance.
A big industry problem is that most of the data at the scale these experiments need is private. We thought that was worth changing, so we've opened this up.
Getting to 100,000 hours
At this scale, the hard part isn't recording more video. It's keeping collection consistent. Small differences in capture setup, environment, task definition, or annotation quality add up fast once you're talking about tens of thousands of hours.
EgoSuite-Open100K was collected by a globally distributed workforce, with tens of thousands of collectors working within a standardized, continuous collection process. We track coverage targets across collector recruitment and geography, the scene library, and task allocation, so growth doesn't come at the cost of consistency.
100,000 hours of the same few environments wouldn't do much for anyone. The full dataset spans:
- 7 environmental categories: home, hospitality, retail, sports, logistics, office, industry
- 128 scene types: bedrooms, kitchens, retail floors, warehouses, assembly lines, offices, and more
- 18 task categories: assembly and installation, cooking, inventory management, tool use, repair and maintenance, packing, and other everyday and professional work
From video to structured behavior
Raw egocentric video carries a lot of signal, but most of it isn't directly usable without structure on top. The collection is organized into two capture configurations, each split into sub-SKUs by annotation depth:
EgoStandard — the bulk of the dataset, standard egocentric capture.
EgoStand: hand pose (80,000 h planned)EgoStand-Body: hand pose + full body pose (10,000 h planned)
EgoPro — adds a wrist-mounted camera for close-range interaction, contact, and grasping that a head-mounted view alone tends to miss (hands leaving frame, occlusion right at the moment of contact, fine detail that's a few pixels at head height).
EgoProStandard: wrist + hand pose (8,000 h planned)EgoProStandard-Body: wrist + hand pose + full body pose (2,000 h planned)
If you want a taste of all of it before committing to the full download, EgoDemo packages 50 hours pulled from all four sub-SKUs above, plus two raw-video variants.
Three annotation types run across the dataset:
Hand pose: Hands are the hard part of egocentric data — small in frame, fast-moving, frequently occluded, often interacting with visually cluttered objects. We built our hand-pose pipeline specifically around these failure modes.
Body pose: Ties hand and arm movement back to the surrounding task and environment.
Event-level semantic annotation: On selected subsets, gives higher-level structure over time — what's happening, not just what's moving.
Annotation and camera coverage vary by subset — check individual dataset cards on the Hub for exact modality coverage.
Formats
It's released in LeRobot v3 so it's training-ready and streamable straight from the Hub, and in MCAP for teams running their own robotics/multimodal data pipelines.
Get started
- Browse the EgoSuite-Open100K collection on Hugging Face →
- EgoStandard →
- EgoPro →
- EgoDemo (50-hour sample) →
What it's for
We released this openly because we want to see what people build with it. Some directions we expect to be useful:
- VLA (vision-language-action) model pretraining
- World model pretraining
- Human-to-robot behavior transfer
- Egocentric representation learning
- Hand-object interaction modeling
- Action and activity recognition
- Task and intent understanding
- Long-horizon activity understanding
- Human and hand pose estimation
- Learning representations of real-world manipulation
We'd also be surprised if that list covers everything people end up doing with it.
A step toward shared standards
Scale isn't the only thing holding egocentric data back — capture conventions, annotation schemas, sensor configs, and storage formats are still fragmented across the field, which makes datasets hard to combine and compare. Through our work with the EgoVerse consortium, we're aligning EgoSuite with emerging standards for egocentric data capture, annotation, and sharing, and hoping this contributes to that conversation as much as it contributes more data.
This is 10,000 of 100,000
We're releasing the rest progressively, and this first batch is also a chance for the community to shape what comes next. If you're training or evaluating on EgoSuite, we'd like to know:
- Which tasks are most useful to you?
- What environments feel underrepresented?
- Which annotations matter most?
- What additional modalities would help?
- Where does the dataset fall short?
Open a discussion on the relevant dataset repo on the Hub and tell us what you find — it'll help shape the next 90,000 hours. You can also join the Lightwheel Discord to ask questions, share what you're building with the dataset, and talk directly with the team.
Licensing
The released subsets are available for academic research and commercial training. For licensing details, modality coverage, and subset-specific structure, see the individual dataset cards on Hugging Face.
Resources
- Hugging Face Collection: EgoSuite-Open100K
- Datasets: EgoStandard · EgoPro · EgoDemo
- Project website: egosuite100k.lightwheel.ai
- Lightwheel on Hugging Face: huggingface.co/LightwheelAI
- Lightwheel: lightwheel.ai
- Discord: discord.gg/29r9Zu4Kk5
This is the first public layer of the data infrastructure we're building for Physical AI, not the finished product. The interesting part starts now — what models learn from it, where it holds up, where it doesn't, and what the community finds once human data at this scale is actually in use.
