humanoidsdata.com

Search

Search datasets, articles, and glossary terms for humanoids and embodied AI.

Browse a curated catalog of datasets for humanoid robots and embodied AI, spanning real-world demonstrations, teleoperation, motion capture, egocentric vision, simulation, manipulation, locomotion, and cross-embodiment robot learning.

83 Datasets · Page 4 of 7

A Franka robot assembling several FurnitureBench furniture models

FurnitureBench

FurnitureBench includes 5,100 successful real-robot teleoperation demonstrations totaling 219.6 hours across nine furniture configurations and three initialization-randomness levels. A single Franka Panda performs long-horizon grasping, reorientation, insertion, and screwing using wrist and front RGB images, detailed proprioception, 8-D actions, rewards, and skill-completion flags. The standardized physical benchmark and matching FurnitureSim environment support imitation learning, offline reinforcement learning, planning, and reproducible household assembly research.

View dataset →
A Franka robot grasping a shaped peg above an FMB assembly board

Functional Manipulation Benchmark (FMB)

FMB provides 22,550 human demonstrations collected with a Franka Panda across single-object and multi-object functional assembly tasks. The robot must grasp procedurally designed parts, reorient or regrasp them using a fixture, and perform precise insertion or sequential assembly under varied geometry, appearance, and initial poses. Four RGB-D views, robot kinematics, end-effector force/torque, object metadata, and camera intrinsics support point-cloud processing, imitation learning, and controlled generalization studies.

View dataset →
A human sharing a workspace with HABIT's two Franka robot arms

HABIT

HABIT contains 10,563 real episodes and 164.19 hours across 60 tasks, with a person sharing the workspace in every demonstration and two teleoperated Franka Research 3 arms performing the robot role. Collaborator, coworker, and supervisor scenarios capture handovers, shared-workspace safety, temporal coordination, and gesture following. Five RGB streams include robot, wrist, human-egocentric, and exocentric views; Cartesian, joint, gripper, language, workflow, and synchronized human/robot subtask annotations support human-aware policy training.

View dataset →
ManipArena robots placing objects into drawers and sorting objects by shape

ManipArena

ManipArena supplies 10,812 teleoperated real-robot trajectories, 13.5 million frames, and about 188 hours across 20 reasoning-oriented tasks, plus paired simulation demonstrations for three tasks. Its fixed bimanual platform handles execution and semantic tasks, while an omnidirectional mobile dual-arm platform handles navigation-conditioned long-horizon tasks. Three RGB views, end-effector and joint signals, currents, mobile-base state, and three levels of language annotation support controlled training and real-robot evaluation.

View dataset →
ABC bimanual robots performing folding, packing, sorting, and assembly tasks

ABC-130K

The ABC-130K paper reports 134,806 real teleoperated episodes totaling 3,553 hours across 195 tasks, while the current official data card lists 130,703 YAM trajectories totaling 3,590.7 hours. Collected on inexpensive dual-arm YAM stations, the data spans pick-and-place, folding, handover, insertion, sorting, tool use, and assembly, including dexterous behaviors such as box folding and extracting cards from wallets. Episodes provide three-camera video, joint and gripper signals, end-effector poses, task instructions, and subtask labels for a 1,552-hour subset, supporting scalable behavior-cloning research.

View dataset →
RoboNet robot arm manipulating objects in a tabletop bin

RoboNet

RoboNet contains roughly 162,000 autonomously collected trajectories and 15 million video frames from seven robot platforms, four institutions, and 113 camera viewpoints. Sawyer, Franka, Baxter, Fetch, Kuka, WidowX, and Google R3 systems interact with hundreds of objects using random exploratory and grasping policies, recording RGB video, end-effector and gripper actions, and corresponding robot state. The open database supports cross-robot visual foresight, inverse models, object relocation, and pretraining for rapid transfer to unseen viewpoints, grippers, environments, and robot hardware.

View dataset →
Franka robot at a RoboSet kitchen manipulation station

RoboSet

RoboSet comprises 100,050 real-robot trajectories: a kitchen collection reporting 30,050 total trajectories, including 9,500 teleoperated demonstrations and additional kinesthetic-playback trajectories, plus 70,000 bin-manipulation trajectories collected through scripts and policies. The kitchen data cover 38 tasks and 12 skills. RoboAgent was trained on a frozen 7,500-trajectory teleoperation subset collected with Franka Emika arms and Robotiq grippers across varied kitchen scenes. The data cover picking, placing, wiping, capping, sliding, object reorientation, and articulated doors or drawers, while language conditioning and automatically generated semantic image augmentations support sample-efficient generalization to unseen objects, tasks, and kitchens.

View dataset →
WidowX robot placing cloth into a toy laundry machine

BridgeData V2

BridgeData V2 currently exposes 60,096 WidowX 250 trajectories across 24 environments and 13 skill families: 50,365 are VR-teleoperated demonstrations and 9,731 are scripted pick-and-place rollouts. Its toy kitchens, tabletops, sinks, and laundry setup cover pick-and-place, pushing, sweeping, doors and drawers, block stacking, cloth folding, and granular media, with natural-language labels, RGB views, limited depth, robot state, and end-effector/gripper actions. It supports goal-image and language-conditioned offline learning and cross-institution generalization.

View dataset →
Human operator teleoperating an RH20T robot arm

RH20T

RH20T provides more than 110,000 contact-rich real-robot manipulation sequences across roughly 147 tasks and seven arm-and-gripper configurations. Haptic teleoperation captured multi-view RGB-D and infrared video, audio, joint and end-effector state, actions, and six-axis force/torque; one configuration also includes fingertip tactile sensing, and each robot sequence has a paired human demonstration video and language description. The dataset targets one-shot imitation and generalization to diverse real-world skills beyond simple pushing and pick-and-place.

View dataset →
Galaxea R1 Lite mobile dual-arm robot

Galaxea Open-World

Galaxea Open-World contains about 100,000 demonstrations and 500 hours of mobile-manipulation behavior across more than 150 tasks, 50 real-world scenes, and over 1,600 objects. Data were teleoperated on the uniform R1-Lite mobile dual-arm platform in residential, catering, retail, and office settings, with head and wrist video, arm, gripper, torso, chassis, end-effector, and IMU state/action streams. Fine-grained bilingual subtask annotations support VLA pretraining, planning, few-shot transfer, and long-horizon whole-body tasks such as table bussing, microwave operation, and bed making.

View dataset →