humanoidsdata.com

Search

Search companies, datasets, articles, and glossary terms for humanoids and embodied AI.

Browse a curated catalog of datasets for humanoid robots and embodied AI, spanning real-world demonstrations, teleoperation, motion capture, egocentric vision, simulation, manipulation, locomotion, and cross-embodiment robot learning.

90 Datasets · Page 5 of 8

RoboNet robot arm manipulating objects in a tabletop bin

RoboNet

RoboNet contains roughly 162,000 autonomously collected trajectories and 15 million video frames from seven robot platforms, four institutions, and 113 camera viewpoints. Sawyer, Franka, Baxter, Fetch, Kuka, WidowX, and Google R3 systems interact with hundreds of objects using random exploratory and grasping policies, recording RGB video, end-effector and gripper actions, and corresponding robot state. The open database supports cross-robot visual foresight, inverse models, object relocation, and pretraining for rapid transfer to unseen viewpoints, grippers, environments, and robot hardware.

View dataset →
Franka robot at a RoboSet kitchen manipulation station

RoboSet

RoboSet comprises 100,050 real-robot trajectories: a kitchen collection reporting 30,050 total trajectories, including 9,500 teleoperated demonstrations and additional kinesthetic-playback trajectories, plus 70,000 bin-manipulation trajectories collected through scripts and policies. The kitchen data cover 38 tasks and 12 skills. RoboAgent was trained on a frozen 7,500-trajectory teleoperation subset collected with Franka Emika arms and Robotiq grippers across varied kitchen scenes. The data cover picking, placing, wiping, capping, sliding, object reorientation, and articulated doors or drawers, while language conditioning and automatically generated semantic image augmentations support sample-efficient generalization to unseen objects, tasks, and kitchens.

View dataset →
WidowX robot placing cloth into a toy laundry machine

BridgeData V2

BridgeData V2 currently exposes 60,096 WidowX 250 trajectories across 24 environments and 13 skill families: 50,365 are VR-teleoperated demonstrations and 9,731 are scripted pick-and-place rollouts. Its toy kitchens, tabletops, sinks, and laundry setup cover pick-and-place, pushing, sweeping, doors and drawers, block stacking, cloth folding, and granular media, with natural-language labels, RGB views, limited depth, robot state, and end-effector/gripper actions. It supports goal-image and language-conditioned offline learning and cross-institution generalization.

View dataset →
Human operator teleoperating an RH20T robot arm

RH20T

RH20T provides more than 110,000 contact-rich real-robot manipulation sequences across roughly 147 tasks and seven arm-and-gripper configurations. Haptic teleoperation captured multi-view RGB-D and infrared video, audio, joint and end-effector state, actions, and six-axis force/torque; one configuration also includes fingertip tactile sensing, and each robot sequence has a paired human demonstration video and language description. The dataset targets one-shot imitation and generalization to diverse real-world skills beyond simple pushing and pick-and-place.

View dataset →
Galaxea R1 Lite mobile dual-arm robot

Galaxea Open-World

Galaxea Open-World contains about 100,000 demonstrations and 500 hours of mobile-manipulation behavior across more than 150 tasks, 50 real-world scenes, and over 1,600 objects. Data were teleoperated on the uniform R1-Lite mobile dual-arm platform in residential, catering, retail, and office settings, with head and wrist video, arm, gripper, torso, chassis, end-effector, and IMU state/action streams. Fine-grained bilingual subtask annotations support VLA pretraining, planning, few-shot transfer, and long-horizon whole-body tasks such as table bussing, microwave operation, and bed making.

View dataset →
DROID Franka Panda robot and camera rig in a lab workspace

DROID

DROID contains 76,000 teleoperated Franka Panda demonstrations, representing 350 hours across 86 tasks and 564 scenes collected by 50 operators at 13 institutions. A standardized portable setup uses two external stereo cameras, a wrist stereo camera, and an Oculus headset, recording images, robot state, end-effector and gripper actions, calibration, and language instructions; updated annotations provide three descriptions for most successful episodes. It supports robust, generalizable manipulation-policy training in homes, offices, and laboratories.

View dataset →
Robot arm manipulating objects in a toy kitchen

Open X-Embodiment

Open X-Embodiment pools more than one million real-robot trajectories from 60 source datasets contributed by 34 laboratories, covering 22 embodiments and 527 skills. The collection spans single arms, bimanual systems, mobile manipulators, and other platforms performing varied manipulation behaviors with household objects, while normalizing observations, actions, and task information into a common RLDS-oriented ecosystem. It was created to study cross-robot pretraining, transfer, and generalist policies such as RT-X.

View dataset →
Two CORE4D human models collaboratively moving a chair

CORE4D

CORE4D combines 1,000 real motion-captured sequences of two people collaboratively rearranging household objects with 10,000 retargeted sequences spanning roughly 3,000 virtual object shapes. Real sequences include SMPL-X human and object meshes, egocentric RGB, allocentric RGB-D, calibrated poses, and 2D segmentation masks across carrying, handover, joining, and obstacle-navigation scenarios. It targets collaborative motion forecasting, interaction synthesis, and transfer of human interaction patterns to robot policies.

View dataset →
HUMOTO human model carrying an object

HUMOTO

HUMOTO provides 735 professionally cleaned motion-capture sequences totaling 7,875 seconds at 30 fps, with interactions involving 63 modeled objects and 72 articulated parts. Its 4D representation covers full-body and hand motion, object geometry, and purposeful multi-object activities ranging from cooking and laptop use to carrying furniture and outdoor picnics. Scene-driven scripts and physically refined motion support motion generation, vision, animation, robotics, and embodied-AI research.

View dataset →