Data & collection
Human–object interaction
Human–object interaction is the physical and semantic relationship between a person and an object while the person observes, reaches, grasps, moves, uses, or otherwise acts on it. In robotics datasets, the term often refers to recordings and annotations that connect human body or hand motion with object identity, pose, contact, action, and task context.
Also known as: HOI, human-object interaction, human object interaction
Updated
Interaction includes motion, contact and meaning
HOI research can ask where the hands and object are, when contact changes, which action is occurring or what function the object serves. A video action label captures only one layer. Richer records can include 3D body pose, hand pose, object pose, meshes, contact and temporal segments.
HOI4D combines egocentric RGB-D with frame-level hand, object, motion and action annotations. GRAB captures full-body motion, object pose and body-object contact during grasping and manipulation. These datasets illustrate why HOI is broader than isolated hand detection.
Hand-object interaction is a narrower case
Hand-object data focuses on the hands, fingers and manipulated item. Whole-body HOI can include torso support, walking, sitting or carrying. Human–robot interaction is a different concept: its second participant is a robot, although an episode can contain both kinds of interaction.
The term also does not imply executable robot actions. A human may use anatomy, touch and force that the target robot does not possess.
Value for robot learning depends on reconstruction
HOI data can teach task semantics, object affordances, likely contact regions and human motion priors. Turning it into policy supervision may require pose estimation, 3D reconstruction, contact inference and retargeting to the robot embodiment.
Dataset documentation should identify viewpoints, participants, objects, capture hardware, coordinate frames, annotation methods, contact definition and permitted uses. Estimated hand or object poses should remain distinguishable from measurements, and interaction labels should state their temporal scope.
Sources
Related terms
Data & collection
Egocentric data
Egocentric data is sensor data recorded from the viewpoint of the person or robot performing an activity, most commonly with a head- or body-mounted camera. It can also include audio, gaze, depth or inertial signals. For humanoid learning, it shows hands, objects and actions from an actor-centred perspective.
Data & collection
Motion capture
Motion capture is the measurement and reconstruction of a person’s or object’s movement over time, commonly as joint positions, orientations or a fitted body model. Optical markers, cameras and inertial sensors can supply the measurements. Humanoid robotics uses the resulting motion sequences for analysis, imitation and retargeting to a robot body.
Data & collection
Demonstration
A demonstration is a recorded example of how an intended task or behaviour is performed, usually represented as a time-aligned sequence of observations, states and actions. For humanoid robot learning, demonstrations may come from teleoperation, kinaesthetic guidance, motion capture or autonomous experts and provide targets for imitation.
Simulation & transfer
Motion retargeting
Motion retargeting is the adaptation of a recorded or generated motion from one body to another with different proportions, joints or limits. For humanoid robots, it maps source poses or trajectories into robot configurations while preserving task-relevant relationships such as contacts and end-effector paths and satisfying kinematic, balance, collision and actuator constraints.
Hardware & control
Robot manipulation
Robot manipulation is a robot's controlled physical interaction with objects or its environment to change or maintain their state. It includes grasping, carrying, pushing, pulling, inserting, wiping, folding, tool use, and other tasks performed through selective contact. Manipulation can use a gripper, hand, tool, arm, or another part of the robot.