humanoidsdata.com

Search

Search companies, datasets, articles, and glossary terms for humanoids and embodied AI.

Browse a curated catalog of datasets for humanoid robots and embodied AI, spanning real-world demonstrations, teleoperation, motion capture, egocentric vision, simulation, manipulation, locomotion, and cross-embodiment robot learning.

90 Datasets · Page 6 of 8

Human and motion-capture avatar performing a bimanual bowl manipulation task

KIT Whole-Body Human Motion Database

The KIT Whole-Body Human Motion Database currently reports 2,972 motion experiments, 238 subjects, 164 objects, and 41.61 hours of C3D motion within a 2.1-TB collection. Motion capture records whole-body human and object trajectories, subject-object relations, video and other sensor files, with manual hierarchical motion tags and a dedicated motion-language subset. Its Master Motor Map normalizes subject-specific kinematics and dynamics, enabling motion analysis, primitive learning, and transfer of manipulation and locomotion behaviors to differently embodied humanoid robots.

View dataset →
Dexterous robot hand grasping a cup in the HRDexDB capture setup

HRDexDB

HRDexDB pairs human grasps with executions by four dexterous robot-hand embodiments on the same 100 objects, totaling 2,100 sequences and 24 million frames. A synchronized rig of 21 exocentric and two egocentric RGB cameras records reconstructed 3D hand motion, robot state, object geometry and 6-DoF trajectories, success labels, and fingertip contact forces where tactile hardware is available. It supports human-to-robot contact transfer, cross-embodiment grasp retrieval, and evaluation of hand and object pose estimation under occlusion.

View dataset →
Egocentric human demonstrations of varied tabletop manipulation tasks

EgoDex

EgoDex contains 829 hours, 338,000 demonstrations, and 90 million 1080p frames across 194 tabletop manipulation tasks. Apple Vision Pro and ARKit record natural bare-hand demonstrations with egocentric RGB, calibrated camera parameters, upper-body pose, 25 tracked joints per hand, confidence values, and natural-language descriptions. Tasks range from tying shoelaces and folding laundry to tightening screws, dealing cards, and inserting batteries; the dataset supports dexterous trajectory prediction, inverse dynamics, robotics pretraining, and egocentric perception research.

View dataset →
UniDex robot hands manipulating varied household objects

UniDex

UniDex-Dataset provides more than 50,000 robot-executable trajectories and 9 million paired image-point-cloud-action frames across eight dexterous hands with 6–24 active degrees of freedom. It derives these trajectories from four egocentric RGB-D human-manipulation datasets using human-in-the-loop fingertip retargeting, contact adjustment, hand masking, and robot-hand rendering. Language-aligned tasks include phone use, opening cartons, cooking with a spatula, coffee making, sweeping, cutting bags, and mouse operation, supporting 3D VLA pretraining and transfer between hand embodiments.

View dataset →
Retargeted dexterous robot hand grasping a cylindrical object with tactile contacts

HORA (Hand–Object to Robot Action Dataset)

HORA contains about 150,000 hand-object trajectories drawn from 63,141 tactile motion-capture demonstrations, 23,560 custom RGB-D recordings, and 66,924 trajectories derived from public datasets. RoboWheel reconstructs hand and object motion, refines contact plausibility, retargets it to robot arms, dexterous hands, and humanoids, and augments executable rollouts in Isaac Sim. Modalities include MANO parameters, object geometry and 6-DoF pose, contact, tactile maps, language task descriptions, robot RGB observations, joint states, and canonical end-effector actions for imitation and VLA training.

View dataset →
Franka robot performing a kitchen manipulation task with MimicPlay

MimicPlay

MimicPlay combines multiview human play with a smaller set of teleoperated, single-arm robot demonstrations for 14 long-horizon manipulation tasks in six real environments. The TFDS release contains 378 episodes and 7.14 GiB, including front and wrist RGB views, language instructions, end-effector and gripper state, joint state, and robot actions. Human-derived 3D hand plans guide a low-level robot controller on tasks involving kitchens, flowers, whiteboards, sandwiches, and cloth, reducing the robot demonstrations needed for imitation learning.

View dataset →
Robot arm using a UMI gripper to arrange a cup outdoors

Universal Manipulation Interface (UMI)

UMI records human demonstrations using one or two portable, camera-equipped handheld parallel grippers rather than robots at the collection site. It captures wrist-view visual context, gripper motion and width, and relative inter-gripper trajectories, with collection reaching 111 demonstrations per hour in the reported cup task. Released experiments cover dynamic object tossing, precise cup placement, bimanual cloth folding, and seven-stage dishwashing, and the resulting policies transfer between UR5e and Franka arms and generalize to new objects and environments.

View dataset →
Robot arms executing household manipulation tasks from EgoInfinity trajectories

EgoInfinity preview

EgoInfinity converts arbitrary-view web videos of human manipulation into metric 4D hand-object representations without manual annotation or wearable capture. The preview includes 106 processed clips, with robot retargeting files for 104 clips across Unitree G1, Robonaut 2, Franka, and XLeRobot embodiments. Outputs include MANO hands, depth, object geometry and 6-DoF pose, interaction states, and executable joint trajectories, supporting video-to-action learning for tasks such as grasping, cutting, wiping, and pouring.

View dataset →
Egocentric demonstrations placing a stapler and serving food

TASTE-Rob

TASTE-Rob contains 100,856 fixed-view, 1080p egocentric hand–object interaction videos—about 9 million frames—each under eight seconds and aligned to one language instruction. Its 75,389 single-hand and 25,467 double-hand clips cover kitchens, bedrooms, dining and office tables, with actions such as picking, placing, pushing, pouring, cleaning, and drawer use. The dataset targets task-oriented hand–object video generation and imitation learning, using scene and object diversity plus precise action-language alignment to improve manipulation generalization.

View dataset →
Egocentric human demonstrations handling bread and a drying rack

EgoLive

EgoLive contains 1,680 hours across 65,866 episodes and 346 task-oriented human routines, recorded in unconstrained settings such as home services, retail, pharmacies, and practical work. A custom JoyEgoCam captures synchronized 2160×2160 stereo RGB at 60 Hz and 200 Hz IMU data. Automated annotations add depth, six-DoF camera motion, 3D hand keypoints, hand and object masks, scene reconstruction, subtask boundaries, and hierarchical instructions for manipulation grounding, workflow decomposition, and downstream human-to-robot transfer.

View dataset →