humanoidsdata.com

Search

Search companies, datasets, articles, and glossary terms for humanoids and embodied AI.

Browse a curated catalog of datasets for humanoid robots and embodied AI, spanning real-world demonstrations, teleoperation, motion capture, egocentric vision, simulation, manipulation, locomotion, and cross-embodiment robot learning.

90 Datasets · Page 1 of 8

First-person Eidon Tracker POV frame showing a hand gripping grey fabric on pink patterned bedding; still extracted at 10 seconds from recording 1000, by Solidic Labs Inc (Eidon AI), CC BY 4.0

Eidon Tracker POV

Eidon Tracker POV contains 13,451 egocentric recordings totalling 1,273.8 hours from 27 contributors performing household tasks, predominantly folding laundry. Head-mounted MP4 video spans 1080p to 4K, mostly at 30 fps. The companion tracker-pov-imu repository provides 24 Hz orientation quaternions for a seven-point harness on the hands, forearms, upper arms and chest, joined to video metadata by recording_id. Raw accelerometer, gyroscope and magnetometer readings are available for 2,841 recordings; 129 recordings have fewer than seven sensor slots. Video and IMU capture are not hardware-synchronised. Metadata includes activity labels and quality-control flags, including invalid recordings. Footage is unredacted; contributor-level splits and quality filtering are recommended. The separate 306-hour video-only bucket is not included in these totals.

View dataset →
Grid of T-Rex bimanual dexterous robot demonstrations across 20 household-object manipulation primitives

T-Rex Dataset

The downloadable T-Rex release contains 5,464 teleoperated episodes—5,473,459 frames, or about 50 hours at 30 Hz—collected on a fixed-base bimanual Dexmate Vega-1 with two Sharpa Wave dexterous hands; the paper reports a 100-hour full corpus. The release spans 207 household objects and 22 motor primitives, including 5,370 language-annotated trajectories. Each episode aligns three RGB views, bimanual joint states and target actions, ten raw fingertip-tactile video streams, ten deformation-map streams, and per-fingertip 6-axis wrench signals in LeRobotDataset v3.0.

View dataset →
First-person EgoSuite-Open100K frame in an industrial workspace with 3D hand-pose skeleton overlays on both hands

EgoSuite-Open100K

EgoSuite-Open100K is a staged 100,000-hour release of egocentric human demonstrations spanning more than 15,000 tasks and 15,000 distinct scenes across 18 task categories and seven environment categories. EgoStandard allocates 90,000 planned hours to head-view video with synchronized left and right 3D hand poses and optional full-body pose; EgoPro allocates 10,000 planned hours and adds wrist-view video. Data is distributed in LeRobot v3 and MCAP, with event-level semantic annotations on selected subsets; EgoDemo provides a 50-hour sample drawn from the four annotated sub-SKUs.

View dataset →
LIBERO-Plus simulated robot scenes showing changes in camera viewpoint, robot initial state, sensor noise, object layout, background texture, and lighting

LIBERO-Plus

LIBERO-Plus expands LIBERO into a robustness benchmark with 10,030 test-only simulated task instances spanning seven perturbation factors and 21 subdimensions: object layout, camera viewpoint, robot initial state, language, lighting, background texture, and sensor noise. Tasks are generated automatically and stratified into five difficulty levels. The associated mixed training release contains more than 20,000 successful trajectories; its official LeRobot package exposes 14,347 episodes, 2,238,036 frames at 20 Hz, and 40 task labels with front and wrist RGB, robot state, and actions.

View dataset →
Mosaic from the RekaDaily-10k release showing daily-life and workplace scenes including cooking, driving, ironing, and electronics work

RekaDaily-10k

RekaDaily-10k is an incrementally released corpus targeting 10,312 hours of unscripted first-person daily-life video recorded by paid collectors with head-mounted and handheld phones in homes and workplaces across multiple regions. The raw tier preserves sessions as collected and provides activity, lighting, duration, frame-rate, resolution, frame-count, codec, and salted collector metadata; Reka also announced a processed tier of short machine-captioned clips. Roughly 1,670 hours of the full corpus are native 4K.

View dataset →
A mosaic of AIRoA MoMa 5k Human Support Robots performing mobile manipulation tasks

AIRoA MoMa 5k

AIRoA MoMa 5k contains 1,184,259 successful primitive-action episodes spanning 5,025.1 hours and 180,905,084 frames of teleoperated mobile manipulation with 44 Toyota Human Support Robots across five sites. The public, train-only LeRobot v3.0 release covers 68 main short-horizon task templates and combines head- and hand-camera RGB video with robot state and actions, end-effector pose, wrist force-torque history, low-level servo telemetry, and hierarchical task metadata.

View dataset →
A mosaic of RealOmni household manipulation demonstrations

10Kh RealOmni-Open Dataset

The 10Kh RealOmni-Open Dataset contains more than 13,000 hours and 5 million clips of bimanual human demonstrations captured across over 10,000 real household scenarios. Collected from more than 3,000 contributors with GenDAS grippers, it covers 30 manipulation skills in 10 scenario groups, including clothing, clutter organization, kitchen cleaning, and shoe handling. The 95 TB MCAP release pairs 1600 × 1296 fisheye video at 30 fps with reconstructed end-effector trajectories, gripper state, 6-axis IMU readings, and tactile-array signals.

View dataset →
A bimanual dexterous robot folding a paper airplane

Robotic Origami Challenge

The Robotic Origami Challenge dataset contains 682 real-world teleoperated demonstrations of a bimanual dexterous robot folding a traditional six-fold paper airplane. Its 51 collection seasons comprise more than 4.7 million frames at 30 fps, with 65-dimensional robot states and actions, joint torques, 10-fingertip force-torque signals, and six synchronized head, wrist, and tactile video streams. The release uses LeRobot v3.0, with many seasons also available in v2.1, for imitation learning, visual-tactile representation learning, and long-horizon deformable-object manipulation.

View dataset →
A mosaic of synchronized HiFi-UMI manipulation demonstrations

HiFi-UMI-2K

HiFi-UMI-2K is a 2,000-hour release of robot-free, bimanual manipulation demonstrations collected with the portable HiFi-UMI system. Each episode includes six microsecond-synchronized camera views, calibrated bimanual trajectories, gripper states, language annotations, subtask boundaries, and per-sample quality metadata. The data is automatically reconstructed and validated through simulation replay, then exported in a LeRobot v3-style Parquet and MP4 format for manipulation-policy pre-training and post-training.

View dataset →