humanoidsdata.com

Search

Search companies, datasets, articles, and glossary terms for humanoids and embodied AI.

Browse a curated catalog of datasets for humanoid robots and embodied AI, spanning real-world demonstrations, teleoperation, motion capture, egocentric vision, simulation, manipulation, locomotion, and cross-embodiment robot learning.

90 Datasets · Page 7 of 8

Egocentric hands assembling a GoPro camera

OpenEgo

OpenEgo consolidates six public egocentric datasets into 1,107 hours, 119.6 million frames, and 344,500 recordings covering 290 manipulation tasks in more than 600 environments. Its kitchens and indoor rooms include cooking, assembly, and daily activities from at least 258 participants. It normalizes both hands to camera-frame MANO-21 joints and adds timestamped, intention-aligned language primitives naming actions, objects, actors, and navigation targets for dexterous imitation learning, world models, and hierarchical vision-language-action training.

View dataset →
Human and robot demonstrations of folding cloth

EgoMimic

The public EgoMimic release is 243 GB and combines Project Aria human demonstrations with teleoperated data from a low-cost bimanual ViperX-arm rig. Human streams include egocentric RGB, SLAM device pose, and 3D hand tracks; robot streams add wrist RGB, end-effector and joint states, and joint actions. Demonstrations cover continuous object-in-bowl, clothes folding, and grocery packing, and support co-training a shared imitation policy that transfers to objects or scenes encountered only in human data.

View dataset →
Robots performing four EgoVerse manipulation tasks

EgoVerse

EgoVerse’s current living release contains 1,362 hours, about 80,000 episodes, from 2,087 people, spanning 1,965 tasks and 240 scenes. Project Aria glasses, phone-mounted cameras, and partner wearables capture human demonstrations; the unified representation includes egocentric video, calibrated six-DoF camera pose, 21-keypoint 3D poses for each hand, task and object metadata, and dense language where available. Standardized tasks include object-in-container, cup-on-saucer, grocery bagging, clothes folding, scooping, and utensil sorting for reproducible human-to-robot transfer.

View dataset →
Kuavo humanoid sorting parts on an assembly line

LET Base / LET Dex

LET Base and LET Dex are continually updated Kuavo 4 Pro full-size humanoid collections covering bipedal and wheeled platforms; Base advertises over 1,000 hours across 31 sub-task scenarios and 117 atomic skills, while Dex emphasizes contact sensing. The releases cover factory sorting and feeding, logistics, hotel services, medical settings, and household tasks using grippers or dexterous hands. Data includes head and wrist RGB-D, joint states and actions, IMU, language step labels, and—in Dex—360-cell fingertip pressure arrays plus six-axis force and torque.

View dataset →
EgoHumanoid robot arms placing a can into a recycling bin

EgoHumanoid sample

The public sample contains 50 Unitree G1 robot-teleoperation episodes and one robot-free human episode—38,905 rows and 160 MB—rather than the full research corpus. Human data is captured with PICO VR, body trackers, and a head-mounted ZED camera; robot data uses VR commands for G1 navigation, wrists, and Dex3 grasping. Both LeRobot-v2 subsets provide egocentric RGB, synchronized whole-body actions, metadata, and language task descriptions for testing human–humanoid co-training and fine-tuning for indoor and outdoor loco-manipulation.

View dataset →
AgiBot G2 performing a household refrigerator task

AgiBot World 2026

AgiBot World 2026 is a 12.8 TB staged real-world release collected on AgiBot G2 robots in commercial, home, and general-purpose environments, and it continues to expand. Long-horizon examples include refrigerated-shelf restocking and cart-to-shelf placement. LeRobot episodes provide head and hand camera streams, robot states and actions, depth-capable sensor configurations, and hierarchical language instructions, key frames, object boxes, success/error/intervention labels, and tactile sensing for applicable end effectors, supporting hierarchical policies and language-grounded manipulation.

View dataset →
Humanoid robot completing a tabletop pouring task

PH2D

PH2D contains 26,824 task-oriented egocentric human demonstrations and 1,552 aligned robot demonstrations, distributed as a 16.2 GB release. Human sequences are captured with Apple Vision Pro or Quest/Vision Pro plus ZED, while robot data comes primarily from Unitree H1 and supports H1-2 transfer. Grasping, placing, cup passing, and pouring episodes pair RGB with 3D head, wrist, and fingertip poses and language instructions, enabling human–humanoid co-training in a shared action space retargeted to robot actions.

View dataset →
HumanPlus humanoid manipulating a shoe while seated

HumanPlus task data

HumanPlus task data consists of real-robot whole-body demonstrations collected by visually shadowing a nearby human on a customized 33-DoF Unitree H1 with two six-DoF hands. The public download contains warehouse unloading, sweatshirt folding, object rearrangement, and robot greeting demonstrations; the paper additionally reports shoe wearing followed by walking and typing, which are not present in the linked release. Released records support behavior cloning from binocular egocentric RGB and humanoid body and hand targets; the broader system learns a retargeted controller from 40 hours of AMASS motion.

View dataset →
DreamDojo robots interacting with objects across varied real-world environments

DreamDojo GR-1 post-training data

The public DreamDojo package provides 74.3 GB and roughly 7.55 million rows of real Fourier GR-1 manipulation data plus evaluation sets for adapting and testing a robot world model. It contains RGB sequences, robot state, relative actions, end-effector information, task descriptions, and coarse and fine human-action annotations for object pickup, placement, transfer, and other interactions. The release is specifically for target-embodiment post-training and evaluation; it does not contain DreamDojo’s separate 44K-hour human-video pretraining corpus.

View dataset →
Unitree G1 humanoid sorting fruit between colored plates

NVIDIA GR00T Teleop G1

NVIDIA’s GR00T Teleop G1 dataset contains 1,000 real teleoperation trajectories of a Unitree G1 with tri-finger hands performing language-conditioned fruit pick-and-place. Four subsets cover apples, pears, grapes, and starfruit, with upper-body control used to move the requested fruit to a target receptacle. Each trajectory includes 640×480 RGB video at 20 FPS plus 43-D full-body-and-hand state and action vectors in MP4 and HDF5, serving as reference data for GR00T fine-tuning.

View dataset →
Leju Kuavo humanoid and VR operator folding clothing

RoboCOIN humanoid subsets

RoboCOIN contains 180K+ real-world bimanual demonstrations covering 421 tasks, 16 residential, commercial, and workplace scenarios, and 15 robot platforms. Its humanoid subsets include the dexterous-hand Unitree G1edu-u3, collected through motion capture or an exoskeleton, and Leju Kuavo 4 Pro, collected through VR teleoperation. Multiview RGB, joint and end-effector state, gripper articulation, and hierarchical trajectory-, subtask-, and frame-level annotations support standardized multi-embodiment bimanual learning.

View dataset →
RoboMIND robots performing cup handling, wiping, drawer, and toaster tasks

RoboMIND 2.0

RoboMIND 2.0 comprises 310K+ real dual-arm trajectories totaling over 1,000 hours across six embodiments; its project page and abstract report 739 tasks, while the paper body and tables report 759. The embodiments include the bipedal Tien Kung humanoid and wheeled Tian Yi platform. Unified teleoperation produced 20K mobile-manipulation and 12K tactile-enhanced episodes, with RGB-D or multiview RGB, proprioception, actions, and fine-grained natural-language annotations. An additional 20K digital-twin simulation trajectories supports sim-to-real research, long-horizon bimanual learning, contact-rich manipulation, and hierarchical VLA training.

View dataset →