humanoidsdata.com

Search

Search datasets, articles, and glossary terms for humanoids and embodied AI.

Browse a curated catalog of datasets for humanoid robots and embodied AI, spanning real-world demonstrations, teleoperation, motion capture, egocentric vision, simulation, manipulation, locomotion, and cross-embodiment robot learning.

83 Datasets · Page 6 of 7

Robot arms executing household manipulation tasks from EgoInfinity trajectories

EgoInfinity preview

EgoInfinity converts arbitrary-view web videos of human manipulation into metric 4D hand-object representations without manual annotation or wearable capture. The preview includes 106 processed clips, with robot retargeting files for 104 clips across Unitree G1, Robonaut 2, Franka, and XLeRobot embodiments. Outputs include MANO hands, depth, object geometry and 6-DoF pose, interaction states, and executable joint trajectories, supporting video-to-action learning for tasks such as grasping, cutting, wiping, and pouring.

View dataset →
Egocentric demonstrations placing a stapler and serving food

TASTE-Rob

TASTE-Rob contains 100,856 fixed-view, 1080p egocentric hand–object interaction videos—about 9 million frames—each under eight seconds and aligned to one language instruction. Its 75,389 single-hand and 25,467 double-hand clips cover kitchens, bedrooms, dining and office tables, with actions such as picking, placing, pushing, pouring, cleaning, and drawer use. The dataset targets task-oriented hand–object video generation and imitation learning, using scene and object diversity plus precise action-language alignment to improve manipulation generalization.

View dataset →
Egocentric human demonstrations handling bread and a drying rack

EgoLive

EgoLive contains 1,680 hours across 65,866 episodes and 346 task-oriented human routines, recorded in unconstrained settings such as home services, retail, pharmacies, and practical work. A custom JoyEgoCam captures synchronized 2160×2160 stereo RGB at 60 Hz and 200 Hz IMU data. Automated annotations add depth, six-DoF camera motion, 3D hand keypoints, hand and object masks, scene reconstruction, subtask boundaries, and hierarchical instructions for manipulation grounding, workflow decomposition, and downstream human-to-robot transfer.

View dataset →
Egocentric hands assembling a GoPro camera

OpenEgo

OpenEgo consolidates six public egocentric datasets into 1,107 hours, 119.6 million frames, and 344,500 recordings covering 290 manipulation tasks in more than 600 environments. Its kitchens and indoor rooms include cooking, assembly, and daily activities from at least 258 participants. It normalizes both hands to camera-frame MANO-21 joints and adds timestamped, intention-aligned language primitives naming actions, objects, actors, and navigation targets for dexterous imitation learning, world models, and hierarchical vision-language-action training.

View dataset →
Human and robot demonstrations of folding cloth

EgoMimic

The public EgoMimic release is 243 GB and combines Project Aria human demonstrations with teleoperated data from a low-cost bimanual ViperX-arm rig. Human streams include egocentric RGB, SLAM device pose, and 3D hand tracks; robot streams add wrist RGB, end-effector and joint states, and joint actions. Demonstrations cover continuous object-in-bowl, clothes folding, and grocery packing, and support co-training a shared imitation policy that transfers to objects or scenes encountered only in human data.

View dataset →
Robots performing four EgoVerse manipulation tasks

EgoVerse

EgoVerse’s current living release contains 1,362 hours, about 80,000 episodes, from 2,087 people, spanning 1,965 tasks and 240 scenes. Project Aria glasses, phone-mounted cameras, and partner wearables capture human demonstrations; the unified representation includes egocentric video, calibrated six-DoF camera pose, 21-keypoint 3D poses for each hand, task and object metadata, and dense language where available. Standardized tasks include object-in-container, cup-on-saucer, grocery bagging, clothes folding, scooping, and utensil sorting for reproducible human-to-robot transfer.

View dataset →
Kuavo humanoid sorting parts on an assembly line

LET Base / LET Dex

LET Base and LET Dex are continually updated Kuavo 4 Pro full-size humanoid collections covering bipedal and wheeled platforms; Base advertises over 1,000 hours across 31 sub-task scenarios and 117 atomic skills, while Dex emphasizes contact sensing. The releases cover factory sorting and feeding, logistics, hotel services, medical settings, and household tasks using grippers or dexterous hands. Data includes head and wrist RGB-D, joint states and actions, IMU, language step labels, and—in Dex—360-cell fingertip pressure arrays plus six-axis force and torque.

View dataset →
EgoHumanoid robot arms placing a can into a recycling bin

EgoHumanoid sample

The public sample contains 50 Unitree G1 robot-teleoperation episodes and one robot-free human episode—38,905 rows and 160 MB—rather than the full research corpus. Human data is captured with PICO VR, body trackers, and a head-mounted ZED camera; robot data uses VR commands for G1 navigation, wrists, and Dex3 grasping. Both LeRobot-v2 subsets provide egocentric RGB, synchronized whole-body actions, metadata, and language task descriptions for testing human–humanoid co-training and fine-tuning for indoor and outdoor loco-manipulation.

View dataset →
AgiBot G2 performing a household refrigerator task

AgiBot World 2026

AgiBot World 2026 is a 12.8 TB staged real-world release collected on AgiBot G2 robots in commercial, home, and general-purpose environments, and it continues to expand. Long-horizon examples include refrigerated-shelf restocking and cart-to-shelf placement. LeRobot episodes provide head and hand camera streams, robot states and actions, depth-capable sensor configurations, and hierarchical language instructions, key frames, object boxes, success/error/intervention labels, and tactile sensing for applicable end effectors, supporting hierarchical policies and language-grounded manipulation.

View dataset →
Humanoid robot completing a tabletop pouring task

PH2D

PH2D contains 26,824 task-oriented egocentric human demonstrations and 1,552 aligned robot demonstrations, distributed as a 16.2 GB release. Human sequences are captured with Apple Vision Pro or Quest/Vision Pro plus ZED, while robot data comes primarily from Unitree H1 and supports H1-2 transfer. Grasping, placing, cup passing, and pouring episodes pair RGB with 3D head, wrist, and fingertip poses and language instructions, enabling human–humanoid co-training in a shared action space retargeted to robot actions.

View dataset →