humanoidsdata.com

Search

Search companies, datasets, articles, and glossary terms for humanoids and embodied AI.

By Lumi · Humanoid robot data · · 9 min read

How InterMimicGen Grows Humanoid Training Data

Can a humanoid tracker expand one human-object motion into many physically executable training trajectories without changing the task? InterMimicGen shows that it can do this in simulation: make small edits to a verified interaction, train the tracker on the proposed variations, execute them under physics, and keep only the rollouts that still complete the intended interaction.

That is more useful than merely perturbing joint angles, but narrower than a robot teaching itself new jobs. The authors’ loop densifies coverage around known demonstrations. It does not invent new task semantics, discover goals in the real world, or prove that every accepted simulation rollout will work on hardware.

A mosaic of humanoid and mobile robots carrying boxes, chairs, luggage, tools, and other objects with InterMimicGen motions

InterMimicGen presents one pipeline across varied human-object interactions and six robot configurations. Source: the authors’ project page and paper.

The loop grows coverage, not task meaning

A captured interaction is one path through a much larger space. A person might lift a box from one position with one stance and one elbow posture. The same task could remain valid with the box farther left, the pelvis lower, the feet wider, or the elbows arranged differently. A policy trained only on the recorded path may never see those nearby solutions.

InterMimicGen treats the interaction type, intended contacts, object, and task outcome as invariants. It then edits either where the interaction happens or how the body realizes it. Object edits translate or rotate the object trajectory. Body edits vary the pelvis, stance, toe angle, or arm posture while holding task-relevant targets fixed.

This is an object-centric representation of the problem. The tracker receives geometry relative to the object, and the retargeter preserves body-object and hand-object relationships instead of copying human joint angles literally. That distinction is essential for loco-manipulation, where moving an object changes balance, contact, reach, and foot placement at the same time.

The closest existing analysis on this site is Humanoid-GPT’s motion-scaling pipeline. Humanoid-GPT studies broad reference-motion tracking after object interactions have been excluded. InterMimicGen asks a different question: can object-coupled motions become new, task-preserving robot references and remain executable as their placement and posture move away from the source?

One dataset record passes through three meanings

The paper begins with human-object motion capture, combining InterAct and the object-interaction portion of HiPHI. The authors report 16,059 motions, 140.68 reference hours, and 157 object entries. Those are human interaction sources, not yet successful robot episodes.

First, motion retargeting maps the source into each robot’s configuration. The coarse stage preserves long-horizon body-object geometry and shapes the grasp. A second cleanup reduces penetration, foot sliding, hand self-collision, joint-limit violations, and contact drift. Sequences the simulator cannot represent—such as a bag hanging from an unmodelled strap—are removed. Robots without articulated fingers also exclude interactions that depend on finger contact.

Second, a reinforcement-learning tracker tries to execute the resulting kinematic reference under simulated dynamics. Its observation includes robot state, future reference frames, and object-relative geometry. Its action is a set of joint targets for proportional–derivative control. A reference that looks plausible geometrically can still fail here by falling, losing the object, penetrating it, or departing too far from the intended motion.

Third, the successful simulation rollout becomes a verified robot motion. This is the most valuable training object in the loop because it records what the controlled embodiment actually did under physics, not only what the retargeter asked it to do.

InterMimicGen pipeline from human-object motion capture through robot retargeting, physics-based tracking, task-preserving edits, and accepted simulation rollouts

Candidate variations enter training before the tracker executes them; only successful, task-preserving rollouts become parents for the next round. Source: Figure 2 of the InterMimicGen paper.

For a dataset buyer, these three meanings should remain separate. A source capture, a retargeted reference, and a simulated execution can all be called a “trajectory”, but they support different claims. The first documents a human solution, the second encodes a robot-specific target, and the third provides evidence that a controller completed an acceptance test.

The tracker is part of the data generator

InterMimicGen is best understood as iterative trajectory augmentation, not a one-shot synthetic-data transform. Each round starts from verified parents, proposes local variations, fine-tunes the tracker, runs the proposals in simulation, and selects successful executions as the next parents.

This combines two ideas. MimicGen showed how object-centric transforms and execution filters could grow robot-manipulation demonstrations from a small seed set. InterMimicGen adds an important complication: the tracker itself must learn to reach the newly proposed conditions. The generation frontier therefore moves only when the policy moves with it.

The method joins retargeting, imitation, execution filtering, and repeated augmentation. Source: the InterMimicGen project page.

The acceptance rule is more concrete than “the motion looks good”. A proposed rollout must reach the reference endpoint without early termination, avoid a sustained fall, pass smoothness checks, and preserve the task outcome. Held-object tasks check terminal hand contact. Placement tasks also check that the object comes to rest on its support. Each candidate is executed with three random seeds; the paper’s appendix requires two successful rollouts for object edits and one for body edits.

Those thresholds are design choices, not universal definitions of valid humanoid data. Passing one of three seeds tolerates more stochastic fragility than passing all three. Completing a supplied reference is also different from responding to an unseen camera observation. A serious release should keep the raw validation results so consumers can choose a stricter filter later.

Why filtering alone stops growing the dataset

The paper’s cleanest ablation freezes the tracker while continuing to generate and filter candidates. In round one, the frozen and fine-tuned trackers accept similar proportions. By round two, the frozen tracker’s reported yield falls to 0.4%, while the fine-tuned tracker accepts 72.9%. In round three the figures are again 0.4% versus 68.9%.

That result supports the authors’ central mechanism: extra sampling cannot keep expanding the frontier if the policy never learns to execute conditions beyond its original competence. The accepted data and the policy form a curriculum. Each verified variation becomes a nearby starting point for the next edit.

Over five rounds, the authors report cumulative reference-set growth from about 20 times the seed set after round one to 150.5 times for Inspire-hand bimanual interactions, 142.0 times for Dex3 bimanual interactions, and 146.4 times for Inspire-hand grasping. On the accepted generated references, the original tracker executed 52.1% to 64.2%, depending on the setting; the evolved trackers executed 98.4% to 98.9%. Both versions retained 100% reported success on the original references.

InterMimicGen examples showing one source interaction expanding into increasingly varied verified robot motions over five rounds

The reported growth is local to each task: later rounds move object placement and body posture farther from verified parents while preserving the source interaction. Source: Figure 6 of the paper.

The growth number should not be read as 150 times more independent human knowledge. All descendants share a source interaction and a hand-designed edit family. The paper also reports that foot sliding rises from roughly 1% of contact frames on original motions to roughly 2% on augmented motions, while joint acceleration and hand jitter increase more moderately. Coverage grows, but the edge of the distribution gets harder and somewhat rougher.

Cross-embodiment transfer is promising but qualitative

The reference construction is evaluated across six configurations: three Unitree G1 hand setups, Booster T1, Booster K1, and the wheeled Dexmate Vega. Real-world examples cover a G1 pulling a suitcase, a G1 with Inspire hands carrying a tripod, a Booster K1 lifting a box, and a Dexmate Vega moving a chair.

A physical G1 executes an augmented box-carry-and-set-down motion. This is qualitative transfer evidence, not a reported hardware success-rate study. Source: the authors’ project page.

This breadth matters because it shows the reference format is not tied to one skeleton or one kind of base. It does not establish a universal policy shared across the bodies. Each platform uses its own configuration and controller. The G1 uses a whole-body low-level controller; the Booster K1 policy uses its own training and deployment stack; the Vega sends time-scaled upper-body trajectories and solves wheeled-base commands separately.

The paper says these low-level controllers track robot motion without object feedback. That makes successful videos evidence that the trajectories are physically coherent enough for those demonstrations. It also limits what they prove: there is no closed-loop visual correction if the physical object starts in a different pose or slips during execution.

What the release does not yet let us reproduce

As of 6 October 2026, the project page links the paper and extensive first-party media, but I found no public InterMimicGen code repository, processed reference collection, training configuration, or generated dataset release. The appendix documents algorithms, metrics, controller settings, and hardware details, which supports technical scrutiny, but an independent team cannot yet rerun the full pipeline from the linked artifacts.

The paper also names two structural limitations. The edit loop cannot create new task semantics; it only broadens realizations of a known interaction. Iterative selection may favor easier objects or compound errors inherited from earlier references. Further limits follow from the evaluation:

  • The headline expansion and success figures are author-reported simulation results, not independent replication.
  • Acceptance tests encode chosen thresholds for contact, falls, smoothness, and terminal object state rather than a complete safety or usefulness standard.
  • Real-robot evidence is qualitative, without a published trial count or success interval across the demonstrated objects and bodies.
  • The loop expands motion references and tracking capability; it does not train perception, task selection, language grounding, or long-horizon recovery.
  • Parent and child motions are correlated, so raw trajectory counts overstate independent behavioural diversity.

This boundary differs from both human motion as training data and synthetic robot-data generation. Human capture supplies the interaction prior. Simulation supplies physical execution and variation. Neither should be hidden behind one undifferentiated “generated data” label.

The useful product is the lineage, not the multiplier

For teams evaluating or packaging InterMimicGen-style data, the important deliverable is a reproducible parent-to-child chain. Each accepted trajectory should preserve:

  • the source dataset, performer, object, licence, capture frame rate, contact labels, and original clip identifier;
  • the target robot model, hand configuration, geometry, joint map, retargeter version, solver settings, and rejected constraints;
  • the parent trajectory and exact object or body edit, including which task invariants were held fixed;
  • simulator, controller, checkpoint, physics randomization, rollout seed, termination reason, and per-check acceptance result;
  • the generated reference separately from the executed state-action rollout and any later real-robot log;
  • dataset-round statistics by task and object, so easier interactions cannot silently dominate later generations.

InterMimicGen’s real contribution is not a magical 150-times multiplier. It is a testable recipe for moving the edge of a humanoid motion distribution: change one known interaction slightly, teach the tracker to reach it, reject failures under physics, and retain successful execution as the next piece of data. That can make sparse demonstrations much more useful. The task catalogue, independent hardware evidence, and reproducible public release still have to grow by other means.