humanoidsdata.com

Search

Search datasets, articles, and glossary terms for humanoids and embodied AI.

Browse a curated catalog of datasets for humanoid robots and embodied AI, spanning real-world demonstrations, teleoperation, motion capture, egocentric vision, simulation, manipulation, locomotion, and cross-embodiment robot learning.

83 Datasets · Page 5 of 7

DROID Franka Panda robot and camera rig in a lab workspace

DROID

DROID contains 76,000 teleoperated Franka Panda demonstrations, representing 350 hours across 86 tasks and 564 scenes collected by 50 operators at 13 institutions. A standardized portable setup uses two external stereo cameras, a wrist stereo camera, and an Oculus headset, recording images, robot state, end-effector and gripper actions, calibration, and language instructions; updated annotations provide three descriptions for most successful episodes. It supports robust, generalizable manipulation-policy training in homes, offices, and laboratories.

View dataset →
Robot arm manipulating objects in a toy kitchen

Open X-Embodiment

Open X-Embodiment pools more than one million real-robot trajectories from 60 source datasets contributed by 34 laboratories, covering 22 embodiments and 527 skills. The collection spans single arms, bimanual systems, mobile manipulators, and other platforms performing varied manipulation behaviors with household objects, while normalizing observations, actions, and task information into a common RLDS-oriented ecosystem. It was created to study cross-robot pretraining, transfer, and generalist policies such as RT-X.

View dataset →
Two CORE4D human models collaboratively moving a chair

CORE4D

CORE4D combines 1,000 real motion-captured sequences of two people collaboratively rearranging household objects with 10,000 retargeted sequences spanning roughly 3,000 virtual object shapes. Real sequences include SMPL-X human and object meshes, egocentric RGB, allocentric RGB-D, calibrated poses, and 2D segmentation masks across carrying, handover, joining, and obstacle-navigation scenarios. It targets collaborative motion forecasting, interaction synthesis, and transfer of human interaction patterns to robot policies.

View dataset →
HUMOTO human model carrying an object

HUMOTO

HUMOTO provides 735 professionally cleaned motion-capture sequences totaling 7,875 seconds at 30 fps, with interactions involving 63 modeled objects and 72 articulated parts. Its 4D representation covers full-body and hand motion, object geometry, and purposeful multi-object activities ranging from cooking and laptop use to carrying furniture and outdoor picnics. Scene-driven scripts and physically refined motion support motion generation, vision, animation, robotics, and embodied-AI research.

View dataset →
Human and motion-capture avatar performing a bimanual bowl manipulation task

KIT Whole-Body Human Motion Database

The KIT Whole-Body Human Motion Database currently reports 2,972 motion experiments, 238 subjects, 164 objects, and 41.61 hours of C3D motion within a 2.1-TB collection. Motion capture records whole-body human and object trajectories, subject-object relations, video and other sensor files, with manual hierarchical motion tags and a dedicated motion-language subset. Its Master Motor Map normalizes subject-specific kinematics and dynamics, enabling motion analysis, primitive learning, and transfer of manipulation and locomotion behaviors to differently embodied humanoid robots.

View dataset →
Dexterous robot hand grasping a cup in the HRDexDB capture setup

HRDexDB

HRDexDB pairs human grasps with executions by four dexterous robot-hand embodiments on the same 100 objects, totaling 2,100 sequences and 24 million frames. A synchronized rig of 21 exocentric and two egocentric RGB cameras records reconstructed 3D hand motion, robot state, object geometry and 6-DoF trajectories, success labels, and fingertip contact forces where tactile hardware is available. It supports human-to-robot contact transfer, cross-embodiment grasp retrieval, and evaluation of hand and object pose estimation under occlusion.

View dataset →
Egocentric human demonstrations of varied tabletop manipulation tasks

EgoDex

EgoDex contains 829 hours, 338,000 demonstrations, and 90 million 1080p frames across 194 tabletop manipulation tasks. Apple Vision Pro and ARKit record natural bare-hand demonstrations with egocentric RGB, calibrated camera parameters, upper-body pose, 25 tracked joints per hand, confidence values, and natural-language descriptions. Tasks range from tying shoelaces and folding laundry to tightening screws, dealing cards, and inserting batteries; the dataset supports dexterous trajectory prediction, inverse dynamics, robotics pretraining, and egocentric perception research.

View dataset →
UniDex robot hands manipulating varied household objects

UniDex

UniDex-Dataset provides more than 50,000 robot-executable trajectories and 9 million paired image-point-cloud-action frames across eight dexterous hands with 6–24 active degrees of freedom. It derives these trajectories from four egocentric RGB-D human-manipulation datasets using human-in-the-loop fingertip retargeting, contact adjustment, hand masking, and robot-hand rendering. Language-aligned tasks include phone use, opening cartons, cooking with a spatula, coffee making, sweeping, cutting bags, and mouse operation, supporting 3D VLA pretraining and transfer between hand embodiments.

View dataset →
Retargeted dexterous robot hand grasping a cylindrical object with tactile contacts

HORA (Hand–Object to Robot Action Dataset)

HORA contains about 150,000 hand-object trajectories drawn from 63,141 tactile motion-capture demonstrations, 23,560 custom RGB-D recordings, and 66,924 trajectories derived from public datasets. RoboWheel reconstructs hand and object motion, refines contact plausibility, retargets it to robot arms, dexterous hands, and humanoids, and augments executable rollouts in Isaac Sim. Modalities include MANO parameters, object geometry and 6-DoF pose, contact, tactile maps, language task descriptions, robot RGB observations, joint states, and canonical end-effector actions for imitation and VLA training.

View dataset →
Franka robot performing a kitchen manipulation task with MimicPlay

MimicPlay

MimicPlay combines multiview human play with a smaller set of teleoperated, single-arm robot demonstrations for 14 long-horizon manipulation tasks in six real environments. The TFDS release contains 378 episodes and 7.14 GiB, including front and wrist RGB views, language instructions, end-effector and gripper state, joint state, and robot actions. Human-derived 3D hand plans guide a low-level robot controller on tasks involving kitchens, flowers, whiteboards, sandwiches, and cloth, reducing the robot demonstrations needed for imitation learning.

View dataset →
Robot arm using a UMI gripper to arrange a cup outdoors

Universal Manipulation Interface (UMI)

UMI records human demonstrations using one or two portable, camera-equipped handheld parallel grippers rather than robots at the collection site. It captures wrist-view visual context, gripper motion and width, and relative inter-gripper trajectories, with collection reaching 111 demonstrations per hour in the reported cup task. Released experiments cover dynamic object tossing, precise cup placement, bimanual cloth folding, and seven-stage dishwashing, and the resulting policies transfer between UR5e and Franka arms and generalize to new objects and environments.

View dataset →