humanoidsdata.com

Search

Search datasets, articles, and glossary terms for humanoids and embodied AI.

Browse a curated catalog of datasets for humanoid robots and embodied AI, spanning real-world demonstrations, teleoperation, motion capture, egocentric vision, simulation, manipulation, locomotion, and cross-embodiment robot learning.

83 Datasets · Page 3 of 7

A simulated Unitree G1 picking a box from a warehouse shelf

Arena G1 Loco-Manipulation

Arena G1 Loco-Manipulation contains five human-teleoperated seeds and 50 automatically generated Unitree G1 trajectories for a simulated box pick-and-place task requiring navigation in Isaac Lab. The MimicGen episodes run at 50 Hz and include 26-DoF desired joint actions and states, timestamps, task-description indices, validity annotations, and 256×256 first-person RGB video. It is intended for behavior cloning, generalist policy post-training, and sim-to-sim or sim-to-real loco-manipulation research.

View dataset →
A simulated Unitree G1 reproducing a captured dance performance

G1 Moves

G1 Moves provides 60 human-performance clips totaling 29.6 minutes, primarily dance and karate, for the 29-DoF Unitree G1. Fifty-nine clips were acquired with markerless LiDAR-and-vision motion capture and one from monocular video, then retargeted into G1 joint trajectories. Each clip includes raw BVH and FBX motion, root and joint trajectories, MuJoCo-derived training states and velocities, validation metadata, and deployable ONNX imitation policies for retargeting, reinforcement learning, visualization, and sim-to-real study.

View dataset →
A Unitree G1 performing agile locomotion and scene-interaction motions

OmniRetarget

OmniRetarget generated more than nine hours of interaction-preserving humanoid trajectories from OMOMO, LAFAN1, and in-house motion capture; the public Unitree G1 release contains four hours because LAFAN1-derived data cannot be redistributed. It covers object carrying, terrain climbing and traversal, and combined object-terrain interaction, with systematic augmentation of object pose, object size, terrain, and embodiment. NPZ files store frame rate, 29-DoF robot joint positions, floating-base pose, and optional object pose for loco-manipulation and reinforcement-learning research.

View dataset →
A Unitree G1 mirroring dynamic motions from human demonstrators

MOSAIC

MOSAIC uses about 64 hours of heterogeneous motion data: 3.1 hours of optical motion capture, 7 hours of inertial capture, 51 hours of public corpora, 2.2 hours of curated GENMO motion, and 1 hour of interface-adaptation data. The release includes AMASS-style human motion, Unitree G1-retargeted NPZ trajectories, raw inertial streams, generation prompts, and roughly 30 minutes each of PICO VR and Noitom data. It supports generalist motion tracking, offline replay, and robust whole-body teleoperation.

View dataset →
A humanoid carrying a box and vacuuming a floor with whole-body control

ALMI-X

ALMI-X contains approximately 81,500 four-second, 200-step Unitree H1-2 trajectories generated by ALMI policies in MuJoCo. It combines AMASS-derived upper-body motions with omnidirectional lower-body velocity commands, including standing, turning, and different movement speeds, and assigns a templated language description to each combination. Files provide robot observations, actions, joint positions, global position, and global orientation for training language-conditioned humanoid locomotion and whole-body control models deployable on real robots.

View dataset →
Reference motions and real Unitree G1 tracking results from PHUMA

PHUMA

PHUMA is a 73-hour corpus of approximately 76,000 physically curated human-motion clips sourced from motion-capture datasets and internet video, then physics-constrained and retargeted for Unitree G1 and H1-2 humanoids. Motions span stationary, angular, vertical, and horizontal locomotion. The linked Hugging Face release is partial: LAFAN1- and LocoMuJoCo-derived motions are excluded and must be obtained and processed through the official repository instructions. Each released file records root translation, root orientation, joint positions, and frame rate, supporting humanoid motion-imitation training, tracking benchmarks, and sim-to-real locomotion policies.

View dataset →
A simulated robot arm completing a multi-stage coffee preparation task

MimicGen

The released MimicGen corpus contains more than 48,000 HDF5 demonstrations across 12 simulated manipulation tasks, including 120 teleoperated human source demonstrations and generated variants spanning reset distributions, objects, and four robot arms. Representative tasks include coffee preparation, threading, mug cleanup, pick-and-place, and multi-part assembly. The data support RGB observations, robot proprioception, and end-effector delta-pose and gripper actions for behavioral cloning and research on scalable automated demonstration generation.

View dataset →
Dexterous humanoid hands preparing to pour from a cup into a bowl

DexMimicGen

DexMimicGen generated 21,000 demonstrations from 60 human source demonstrations across nine MuJoCo tasks and three embodiments: dual Panda arms with grippers, dual Panda arms with dexterous hands, and a GR-1 humanoid with dexterous hands. Source demonstrations were teleoperated using iPhone or Apple Vision Pro interfaces. Tasks include threading, assembly, pouring, tray lifting, drawer cleanup, transport, and can sorting, with observations and coordinated end-effector and hand actions intended for image-based imitation learning and real-to-sim-to-real transfer.

View dataset →
A Panda robot arm working at a toaster in a kitchen

RoboCasa / RoboCasa365

RoboCasa365 defines 365 kitchen tasks and provides more than 2,200 hours of simulated demonstrations: 482 hours of human teleoperation across 300 pretraining tasks and 2,500 pretraining scenes, 1,615 hours of MimicGen data across 60 atomic tasks, and 193 hours of target-task teleoperation across 50 tasks and 10 held-out scenes. The LeRobot-format release includes third-person and eye-in-hand video, proprioception, end-effector, gripper and base actions, language instructions, and per-frame subtask labels for target composite tasks. It supports generalist-policy training and benchmarking in realistic household kitchens.

View dataset →
A Unitree G1 humanoid approaching a table to pick up an object

GRAIL

The current GRAIL release contains 22,312 physics-validated Unitree G1 motions and 5,576,500 frames spanning tabletop and ground pickup, sitting, slopes, curbs, and stair traversal. A fully digital pipeline generates synthetic human-object videos, reconstructs SMPL-X and object 6-DoF motion, retargets it to the humanoid, and validates execution through an Isaac Lab tracking policy. Each sequence includes source video, 4D reconstruction, robot and object trajectories, metadata, and textured USD assets for whole-body policy training and sim-to-real research.

View dataset →
Unitree G1, AgiBot G1, and YAM robots performing manipulation tasks

NVIDIA GR00T X-Embodiment Sim

NVIDIA GR00T X-Embodiment Sim is a simulated post-training collection spanning humanoids, single Panda arms, bimanual Panda variants, grippers, and dexterous hands. Its largest portion contains 240,000 Fourier GR1 humanoid trajectories across 24 tabletop task families; additional releases include 9,000 cross-embodiment bimanual trajectories, 72,000 Panda kitchen trajectories, a 24,000-trajectory downsampled GR1 set, and Unitree G1 loco-manipulation. Vision, language, robot state, and motor actions support cross-embodiment GR00T adaptation.

View dataset →
RT-1 mobile robot completing tabletop manipulation tasks in varied kitchens

RT-1 Robot Action Dataset

The public fractal20220817_data release contains 87,212 real episodes from the larger RT-1 collection, which was gathered by remotely teleoperating a fleet of Everyday Robots mobile manipulators in kitchen-like environments. Episodes cover hundreds of language-described skills such as picking, placing, opening drawers, retrieving objects, pulling napkins, and opening jars. They include RGB, language and embeddings, end-effector and gripper state, success metadata, and arm, gripper, base, and termination actions for multi-task robot learning.

View dataset →