humanoidsdata.com

Search

Search datasets, articles, and glossary terms for humanoids and embodied AI.

← All glossary terms

Data & collection

Human–object interaction

Human–object interaction is the physical and semantic relationship between a person and an object while the person observes, reaches, grasps, moves, uses, or otherwise acts on it. In robotics datasets, the term often refers to recordings and annotations that connect human body or hand motion with object identity, pose, contact, action, and task context.

Also known as: HOI, human-object interaction, human object interaction

Updated

Interaction includes motion, contact and meaning

HOI research can ask where the hands and object are, when contact changes, which action is occurring or what function the object serves. A video action label captures only one layer. Richer records can include 3D body pose, hand pose, object pose, meshes, contact and temporal segments.

HOI4D combines egocentric RGB-D with frame-level hand, object, motion and action annotations. GRAB captures full-body motion, object pose and body-object contact during grasping and manipulation. These datasets illustrate why HOI is broader than isolated hand detection.

Hand-object interaction is a narrower case

Hand-object data focuses on the hands, fingers and manipulated item. Whole-body HOI can include torso support, walking, sitting or carrying. Human–robot interaction is a different concept: its second participant is a robot, although an episode can contain both kinds of interaction.

The term also does not imply executable robot actions. A human may use anatomy, touch and force that the target robot does not possess.

Value for robot learning depends on reconstruction

HOI data can teach task semantics, object affordances, likely contact regions and human motion priors. Turning it into policy supervision may require pose estimation, 3D reconstruction, contact inference and retargeting to the robot embodiment.

Dataset documentation should identify viewpoints, participants, objects, capture hardware, coordinate frames, annotation methods, contact definition and permitted uses. Estimated hand or object poses should remain distinguishable from measurements, and interaction labels should state their temporal scope.

Sources