By Lumi · Humanoid robot data · · 9 min read
Why Humanoid Policies Fail on the Second Object
How can one humanoid policy carry several objects in sequence without forgetting how it handled the first? Humanoid Horizon answers by changing the training distribution: optimize every transport stage together, start later stages from states that earlier stages actually produced, and stop rewarding a rollout when it disturbs an object that was already placed.
The authors report 80.4% full completion for two-object episodes across 350 benchmark scenes and 78.0% across 66 unseen scenes. That is evidence that the three mechanisms work together in this simulator and benchmark. It is not a hardware result: every reported experiment, including the vision-language-action extension, runs in simulation.

Each row is presented as one uninterrupted episode rather than isolated clips joined by resets. Source: the authors’ Humanoid Horizon project page.
The hard data lives between tasks
A single fetch-carry-place cycle can begin from a clean pose on open floor. The second cycle cannot. The humanoid may finish the first placement turned sideways, close to furniture, carrying residual momentum, or with its feet in an awkward stance. It must release the object without moving it, reorient, avoid the new obstacle, and begin the next instruction from that state.
That handoff is why adding tasks is not the same as concatenating demonstrations. The original LHM-Humanoid work framed the transition as a recoverability problem: the first cycle must end in a state from which the next one can begin. Humanoid Horizon asks how to learn those transitions in one policy without a separately trained controller for each stage.
The distinction also separates this work from InterMimicGen. InterMimicGen expands the realizations of a known human-object interaction and filters them through physics. Humanoid Horizon studies a different failure: errors and state shifts that accumulate when several individually familiar interactions must run as one long-horizon task.
For a training-data team, the lesson is concrete. A dataset of successful single-object episodes can be large yet omit the exact states that decide multi-object success: release, recovery, reorientation, approach from a non-canonical pose, and preservation of an earlier result.
Parallel stages fight bias and forgetting
Humanoid Horizon first trains single-object transport, then divides its parallel simulation environments into stage streams. One stream practises the first object, another the second, and so on; all update the same policy. This makes the later stages visible throughout training instead of waiting for the current policy to reach them through a long rollout.
The paper contrasts this with a sequential curriculum learning baseline. That baseline reports 88.3% success on the first transport but only 47.2% full two-object completion across the 350 scenes. The authors call this easy-reward bias: early stages provide frequent learning signal while later transitions remain comparatively rare.
Training a later controller from post-task states creates the opposite risk. Catastrophic forgetting means that updates for the new distribution damage behaviour already learned for the old one. In this paper’s specific setting, a policy adapted to the second transport can lose the ability to repeat the first transport-and-release cycle. Keeping every stage stream active supplies gradients for both distributions.

The shared policy receives simultaneous stage updates; terminal states flow into downstream start buffers; reward gating protects the previous placement. Source: the authors’ method figure and project page.
Parallel training is not simply “more data”. It changes how often the optimizer sees each part of the task and prevents successful early behaviour from monopolising the learning signal. Any release of this kind should therefore report stage sampling, environment allocation, checkpoint origin, and per-stage learning curves—not only the number of simulator steps.
Dynamic starting turns endings into training starts
Parallel streams still need believable initial states. A hand-authored second-stage pose covers only an ideal handoff. It misses the spread of orientations, foot placements, velocities, and object offsets created by the first-stage policy.
Dynamic Starting addresses that mismatch with a small but important feedback loop. At the end of an epoch, each scene’s terminal state for stage one overwrites that scene’s stage-two start buffer. The next epoch trains stage two from the new state. Repeated overwrites expose the policy to a changing sequence of handoffs as upstream behaviour improves.
This is online data generation, but not a steadily growing archive: the implementation stores one current state per scene-stage buffer. Reproducibility therefore depends on logging what would otherwise disappear. Useful lineage includes the upstream policy checkpoint, scene and object IDs, terminal robot and object state, success flags, random seed, epoch, downstream rollout, and replacement history.
The authors’ ablation supports the mechanism within the benchmark. Removing Dynamic Starting lowers full two-object completion from 80.4% to 67.7% on the 350 scenes, and from 78.0% to 63.5% on the 66 unseen scenes. The result is consistent with the handoff-distribution explanation, though it does not isolate which terminal-state dimensions matter most.
A bedroom example shows the policy navigating around furniture before the next manipulation. This is first-party simulation footage, not physical-robot evidence. Source: the project page.
Reward gating changes what counts as success
Per-stage success can hide collateral damage. A robot might place object one correctly, bump it out of position while walking to object two, and still receive credit for completing the second transport. Humanoid Horizon instead defines full-episode success only when all objects reach their goals and remain there.
During a later stage, the method monitors the immediately preceding object. If that object moves beyond a tolerance from its goal, the remaining reward for the rollout is set to zero. Because the same policy repeats this rule at every stage, the authors argue that preservation propagates through the sequence without a separate constraint for every earlier object.
This is a training signal, not a safety shield. It does not prevent contact, roll back the state, or prove collision-free execution. It makes destructive continuation unprofitable under the selected threshold. A reusable dataset should keep both the gate event and the underlying displacement so another team can test a different tolerance or inspect near misses.
Removing reward gating reduces full completion to 69.6% on the 350 scenes and 66.4% on the unseen set. The paper’s metric counts a placement when the object root finishes within 0.2 metres of its target, so those percentages should not be read as millimetre-accurate placement or proof of object stability outside the simulated physics configuration.
Five objects reveal the remaining horizon gap
The most useful result is not that one policy completes a showcase. It is how performance changes as the task gets longer. The authors report the following point estimates:
| Evaluation | Stage 1 | Stage 2 | Stage 3 | Stage 4 | Stage 5 | All required stages |
|---|---|---|---|---|---|---|
| Two objects, 350 benchmark scenes | 96.6% | 81.3% | — | — | — | 80.4% |
| Two objects, 66 unseen scenes | 97.1% | 79.3% | — | — | — | 78.0% |
| Five-object setting | 96.5% | 85.2% | 82.2% | 75.4% | 66.5% | 60.2% |
The in-distribution results come from Table 1, the unseen-scene results from Table 2, and the longer sequence from Table 4. Success declines rather than collapsing, but a 60.2% five-object completion rate still means roughly four in ten evaluated episodes fail under the paper’s criterion.
The paper identifies two recurring failure modes. The first-to-second transition remains unusually difficult because it moves from a canonical start into clutter. The policy can also freeze after reaching an object without securing a stable grasp. The authors attribute part of that failure to equal scene allocation: difficult layouts receive the same training budget as easy ones.
A distinct warehouse sequence places the same training problem around shelving and crates. Source: the authors’ official video grid.
The VLA student removes privileged state, not simulation
The reinforcement-learning teacher observes privileged simulator state and outputs 28 joint position targets for proportional–derivative control. For the paper’s deployment-oriented extension, the authors use DAgger to train separate students from egocentric RGB or depth, body proprioception, a language instruction, and the previous action. A new instruction is supplied after each transport stage.
That is a meaningful observation shift. Student rollouts generate their own states, while the teacher labels corrective actions for those states; this reduces the compounding error that one-pass behaviour cloning can suffer over long sequences. The paper reports that the distilled students retain more of the teacher’s behaviour than students distilled from weaker baselines.
It still does not close the sim-to-real gap. The student images come from the simulator, the instructions arrive at known stage boundaries, and the paper reports no physical humanoid trials. Perception of real clutter, contact uncertainty, actuator limits, latency, and autonomous recognition that one instruction has finished remain untested.
What long-horizon training data should preserve
Humanoid Horizon makes a persuasive case that transitions deserve first-class treatment. A useful long-horizon release would let a downstream team reconstruct more than the final success rate:
- Preserve uninterrupted episodes as well as stage slices, with timestamps linking observations, actions, robot state, object state, goals, rewards, contacts, and termination reasons.
- Mark stage boundaries, but retain the pre- and post-boundary window so release, recovery, and reorientation are not cut away.
- Version every start-state buffer and link it to the upstream rollout and policy checkpoint that produced it.
- Store gate thresholds, raw object displacement, gate activations, and previously completed goals rather than only the shaped reward.
- Report stage-wise and full-episode success across held-out scenes, object combinations, starting states, and longer sequences.
- Separate privileged teacher observations from student-visible inputs and preserve the teacher labels, student-visited states, and DAgger mixing schedule.
The deeper point is that a long task is not merely more frames. Its later actions are conditioned on a history of imperfect outcomes. Humanoid Horizon’s strongest contribution is to make that history part of training through balanced stage updates and rollout-derived handoffs. The reported results show that this can extend simulated task horizons; physical humanoid evidence is the next missing test.