humanoidsdata.com

Search

Search companies, datasets, articles, and glossary terms for humanoids and embodied AI.

Browse a curated catalog of datasets for humanoid robots and embodied AI, spanning real-world demonstrations, teleoperation, motion capture, egocentric vision, simulation, manipulation, locomotion, and cross-embodiment robot learning.

90 Datasets · Page 4 of 8

Reference motions and real Unitree G1 tracking results from PHUMA

PHUMA

PHUMA is a 73-hour corpus of approximately 76,000 physically curated human-motion clips sourced from motion-capture datasets and internet video, then physics-constrained and retargeted for Unitree G1 and H1-2 humanoids. Motions span stationary, angular, vertical, and horizontal locomotion. The linked Hugging Face release is partial: LAFAN1- and LocoMuJoCo-derived motions are excluded and must be obtained and processed through the official repository instructions. Each released file records root translation, root orientation, joint positions, and frame rate, supporting humanoid motion-imitation training, tracking benchmarks, and sim-to-real locomotion policies.

View dataset →
A simulated robot arm completing a multi-stage coffee preparation task

MimicGen

The released MimicGen corpus contains more than 48,000 HDF5 demonstrations across 12 simulated manipulation tasks, including 120 teleoperated human source demonstrations and generated variants spanning reset distributions, objects, and four robot arms. Representative tasks include coffee preparation, threading, mug cleanup, pick-and-place, and multi-part assembly. The data support RGB observations, robot proprioception, and end-effector delta-pose and gripper actions for behavioral cloning and research on scalable automated demonstration generation.

View dataset →
Dexterous humanoid hands preparing to pour from a cup into a bowl

DexMimicGen

DexMimicGen generated 21,000 demonstrations from 60 human source demonstrations across nine MuJoCo tasks and three embodiments: dual Panda arms with grippers, dual Panda arms with dexterous hands, and a GR-1 humanoid with dexterous hands. Source demonstrations were teleoperated using iPhone or Apple Vision Pro interfaces. Tasks include threading, assembly, pouring, tray lifting, drawer cleanup, transport, and can sorting, with observations and coordinated end-effector and hand actions intended for image-based imitation learning and real-to-sim-to-real transfer.

View dataset →
A Panda robot arm working at a toaster in a kitchen

RoboCasa / RoboCasa365

RoboCasa365 defines 365 kitchen tasks and provides more than 2,200 hours of simulated demonstrations: 482 hours of human teleoperation across 300 pretraining tasks and 2,500 pretraining scenes, 1,615 hours of MimicGen data across 60 atomic tasks, and 193 hours of target-task teleoperation across 50 tasks and 10 held-out scenes. The LeRobot-format release includes third-person and eye-in-hand video, proprioception, end-effector, gripper and base actions, language instructions, and per-frame subtask labels for target composite tasks. It supports generalist-policy training and benchmarking in realistic household kitchens.

View dataset →
A Unitree G1 humanoid approaching a table to pick up an object

GRAIL

The current GRAIL release contains 22,312 physics-validated Unitree G1 motions and 5,576,500 frames spanning tabletop and ground pickup, sitting, slopes, curbs, and stair traversal. A fully digital pipeline generates synthetic human-object videos, reconstructs SMPL-X and object 6-DoF motion, retargets it to the humanoid, and validates execution through an Isaac Lab tracking policy. Each sequence includes source video, 4D reconstruction, robot and object trajectories, metadata, and textured USD assets for whole-body policy training and sim-to-real research.

View dataset →
Unitree G1, AgiBot G1, and YAM robots performing manipulation tasks

NVIDIA GR00T X-Embodiment Sim

NVIDIA GR00T X-Embodiment Sim is a simulated post-training collection spanning humanoids, single Panda arms, bimanual Panda variants, grippers, and dexterous hands. Its largest portion contains 240,000 Fourier GR1 humanoid trajectories across 24 tabletop task families; additional releases include 9,000 cross-embodiment bimanual trajectories, 72,000 Panda kitchen trajectories, a 24,000-trajectory downsampled GR1 set, and Unitree G1 loco-manipulation. Vision, language, robot state, and motor actions support cross-embodiment GR00T adaptation.

View dataset →
RT-1 mobile robot completing tabletop manipulation tasks in varied kitchens

RT-1 Robot Action Dataset

The public fractal20220817_data release contains 87,212 real episodes from the larger RT-1 collection, which was gathered by remotely teleoperating a fleet of Everyday Robots mobile manipulators in kitchen-like environments. Episodes cover hundreds of language-described skills such as picking, placing, opening drawers, retrieving objects, pulling napkins, and opening jars. They include RGB, language and embeddings, end-effector and gripper state, success metadata, and arm, gripper, base, and termination actions for multi-task robot learning.

View dataset →
A Franka robot assembling several FurnitureBench furniture models

FurnitureBench

FurnitureBench includes 5,100 successful real-robot teleoperation demonstrations totaling 219.6 hours across nine furniture configurations and three initialization-randomness levels. A single Franka Panda performs long-horizon grasping, reorientation, insertion, and screwing using wrist and front RGB images, detailed proprioception, 8-D actions, rewards, and skill-completion flags. The standardized physical benchmark and matching FurnitureSim environment support imitation learning, offline reinforcement learning, planning, and reproducible household assembly research.

View dataset →
A Franka robot grasping a shaped peg above an FMB assembly board

Functional Manipulation Benchmark (FMB)

FMB provides 22,550 human demonstrations collected with a Franka Panda across single-object and multi-object functional assembly tasks. The robot must grasp procedurally designed parts, reorient or regrasp them using a fixture, and perform precise insertion or sequential assembly under varied geometry, appearance, and initial poses. Four RGB-D views, robot kinematics, end-effector force/torque, object metadata, and camera intrinsics support point-cloud processing, imitation learning, and controlled generalization studies.

View dataset →
A human sharing a workspace with HABIT's two Franka robot arms

HABIT

HABIT contains 10,563 real episodes and 164.19 hours across 60 tasks, with a person sharing the workspace in every demonstration and two teleoperated Franka Research 3 arms performing the robot role. Collaborator, coworker, and supervisor scenarios capture handovers, shared-workspace safety, temporal coordination, and gesture following. Five RGB streams include robot, wrist, human-egocentric, and exocentric views; Cartesian, joint, gripper, language, workflow, and synchronized human/robot subtask annotations support human-aware policy training.

View dataset →
ManipArena robots placing objects into drawers and sorting objects by shape

ManipArena

ManipArena supplies 10,812 teleoperated real-robot trajectories, 13.5 million frames, and about 188 hours across 20 reasoning-oriented tasks, plus paired simulation demonstrations for three tasks. Its fixed bimanual platform handles execution and semantic tasks, while an omnidirectional mobile dual-arm platform handles navigation-conditioned long-horizon tasks. Three RGB views, end-effector and joint signals, currents, mobile-base state, and three levels of language annotation support controlled training and real-robot evaluation.

View dataset →
ABC bimanual robots performing folding, packing, sorting, and assembly tasks

ABC-130K

The ABC-130K paper reports 134,806 real teleoperated episodes totaling 3,553 hours across 195 tasks, while the current official data card lists 130,703 YAM trajectories totaling 3,590.7 hours. Collected on inexpensive dual-arm YAM stations, the data spans pick-and-place, folding, handover, insertion, sorting, tool use, and assembly, including dexterous behaviors such as box folding and extracting cards from wallets. Episodes provide three-camera video, joint and gripper signals, end-effector poses, task instructions, and subtask labels for a 1,552-hour subset, supporting scalable behavior-cloning research.

View dataset →