Models & learning
Object-centric representation
An object-centric representation describes observations, geometry, actions, or trajectories relative to a task-relevant object rather than only in a fixed world frame or robot base frame. In robot learning, it can preserve relationships such as hand-to-object pose and contact while an object is moved, rotated, or transferred between scenes.
Also known as: object-centred representation, object-relative representation
Updated
The object becomes the reference frame
World coordinates describe where something is in a room. Robot-base coordinates describe it relative to the machine. An object-centric representation instead describes task-relevant quantities relative to the object itself: a wrist pose around a handle, a palm-to-surface vector, an intended contact region, or an end-effector path expressed in the object’s frame.
This can make a demonstration easier to reuse. If a cup moves across a table, the world-frame hand trajectory changes, but the relationship between the hand and cup can remain similar. MimicGen uses object-centric segments to transform source demonstrations into new scene configurations, then executes and filters the generated attempts.
Why humanoids need more than an object pose
An object-relative hand target does not determine the whole body. A humanoid still needs feasible feet, balance, collision clearance, joint motion, and contact forces. Large objects can also constrain the torso and both hands over long horizons.
InterMimicGen therefore combines object-relative features with whole-body reference motion and contact intent. Its retargeter preserves a mesh of body-object relationships, while its tracker observes distances from robot links to the object surface. The object-centric part makes translation and rotation edits meaningful; physics-based execution checks whether the resulting whole-body motion is actually feasible.
What the dataset should preserve
“Object-centric” is not a complete data specification. A dataset should state which object frame is used, how that frame was estimated, whether the geometry is a mesh, point cloud or primitive, what scale and units apply, and how symmetries are handled. It should also retain the world and robot frames needed to reconstruct the original scene.
For contact-rich data, preserve the object identifier and revision, geometry source, pose uncertainty, hand and body landmarks, intended and measured contacts, coordinate transforms, timestamps, and any task outcome. Without those fields, an object-relative feature may look portable while hiding errors in scale, calibration, or contact.
Sources
Related terms
Hardware & control
Coordinate frame
A coordinate frame is a defined origin and set of oriented axes used to express positions, orientations, motions, forces, or other spatial quantities. A value has no complete geometric meaning until its frame and convention are known. Transformations relate measurements expressed in frames such as world, robot base, camera, end effector, object, or sensor.
Data & collection
Human–object interaction
Human–object interaction is the physical and semantic relationship between a person and an object while the person observes, reaches, grasps, moves, uses, or otherwise acts on it. In robotics datasets, the term often refers to recordings and annotations that connect human body or hand motion with object identity, pose, contact, action, and task context.
Simulation & transfer
Motion retargeting
Motion retargeting is the adaptation of a recorded or generated motion from one body to another with different proportions, joints or limits. For humanoid robots, it maps source poses or trajectories into robot configurations while preserving task-relevant relationships such as contacts and end-effector paths and satisfying kinematic, balance, collision and actuator constraints.
Models & learning
Pose estimation
Pose estimation is the process of inferring the position and orientation of a body, object, camera, hand, or robot relative to a specified coordinate frame. In three-dimensional robotics this is often called 6D or 6-DoF pose estimation because the result has three translational and three rotational degrees of freedom, even when orientation is stored with more than three numbers.