Models & learning
Catastrophic forgetting
Catastrophic forgetting is a sharp loss of performance on previously learned tasks or data distributions after a neural model is updated on new ones. The new optimization changes shared parameters in ways that overwrite earlier capabilities, so improvement on the latest stage can coexist with regression on skills the model had already acquired.
Also known as: catastrophic interference
Updated
New learning can overwrite old behaviour
Neural networks reuse parameters across examples and tasks. Updating those parameters for a new distribution can move them away from a solution that worked for the old one. The result is more severe than ordinary noise when performance on an earlier capability falls sharply after later training.
The continual-learning study Overcoming catastrophic forgetting in neural networks describes this stability–plasticity problem and evaluates a regularisation method that protects parameters important to earlier tasks. Other approaches include replaying earlier experience, training tasks jointly, allocating task-specific capacity, or distilling prior behaviour. Each method makes different assumptions about which old data, models, or task identities remain available.
Sequential robot training creates a concrete risk
In robot learning, a policy might first learn to approach and carry an object, then be fine-tuned from post-placement states for the next transport. If later-stage updates dominate, the shared controller can improve at continuation while losing the earlier approach or release behaviour.
Humanoid Horizon uses this meaning for multi-object humanoid training. Its parallel stage streams keep learning signal from earlier and later stages active at the same time. That design is evidence against forgetting in the reported benchmark, not a general guarantee that every retained skill will survive new tasks.
Measure retention directly
A final average can hide regression. Training logs should keep per-task or per-stage metrics before and after each curriculum change, the data mixture seen during every update, checkpoint lineage, and evaluation on fixed held-out states.
Retention tests should separate forgetting from distribution shift. A policy failing on a new handoff state may never have learned that state; a policy failing the original stage under its original conditions has lost prior competence. Both matter for long-horizon deployment, but they call for different fixes.
Sources
Related terms
Models & learning
Policy
A policy is the decision rule that maps a robot’s current observations or estimated state, and sometimes a task instruction, to an action or probability distribution over actions. It can be hand-designed or learned from demonstrations, rewards or both. In humanoid robotics, its outputs may be joint targets, torques, end-effector changes or higher-level skills.
Models & learning
Policy distillation
Policy distillation transfers behaviour from one or more teacher policies into a student policy by training the student to match teacher outputs on sampled states. In robotics it can consolidate specialist controllers into one deployable policy, but the student remains limited by the states it visits and the quality and coverage of its teachers.
Models & learning
Transfer learning
Transfer learning uses knowledge learned from one or more source tasks or domains to improve learning or performance on a different target task or domain. In robotics, transfer may reuse representations, model weights, skills, data, or policies across robots, tasks, environments, sensing configurations, or simulation and reality.
Models & learning
Long-horizon task
A long-horizon task is a temporally extended robot task whose success depends on maintaining reliable behaviour across many actions, phases, or dependent subtasks. The term has no universal step-count threshold: it usually signals sequential dependencies, accumulating execution error, delayed outcomes, changing object state, or information that must be remembered beyond the current observation.