Models & learning
Curriculum learning
Curriculum learning is a training strategy that changes which examples, tasks, or initial states a model encounters over time, often moving from easier conditions towards harder ones or expanding the distribution as competence improves. In robot learning, the curriculum controls exploration and learning signal; a poorly balanced sequence can undertrain later stages or degrade earlier skills.
Also known as: curriculum training, training curriculum
Updated
A curriculum changes the training distribution
A fixed training distribution samples the same mixture throughout learning. A curriculum deliberately changes that mixture. It may introduce harder tasks later, widen a range of initial states, increase the length of an episode, or sample conditions near a goal before moving farther away.
Reverse Curriculum Generation demonstrates the last approach for reinforcement learning: starts are sampled near a known goal and expanded outwards as the policy succeeds. The useful abstraction is not “easy examples first” by itself. The curriculum specifies which experience becomes available, when it appears, and how progress changes the sampling rule.
Robot curricula can hide missing stages
Curricula help when rewards are too sparse for a robot to discover a complete task from scratch. They can also bias training. If early states deliver reward frequently, the policy may spend most of its updates on them while later transitions receive little useful signal. Moving entirely to later states can create a different problem if new updates damage earlier competence.
Humanoid Horizon reports this imbalance in sequential object transport. Its alternative keeps stage streams active in parallel and updates one shared policy from all of them, while rollout-derived terminal states change the start distribution for downstream stages.
Record the curriculum as data provenance
A robot dataset produced during training should preserve the curriculum state alongside each episode: task or stage, difficulty parameters, initial-state source, sampling probability, generator or policy checkpoint, success rule, and the condition that advanced or reset the curriculum.
Without that record, a large trajectory set can conceal that most episodes came from easy starts or that harder stages were sampled only after a policy became specialised. Evaluation should use a separately declared distribution so success is not measured only on states chosen by the training curriculum.
Sources
Related terms
Models & learning
Reinforcement learning
Reinforcement learning is a method in which an agent learns a policy by interacting with an environment and optimising cumulative reward. In humanoid robotics, actions change the robot and world, while observations, rewards and episode endings provide experience for improving balance, locomotion or manipulation behaviour.
Models & learning
Long-horizon task
A long-horizon task is a temporally extended robot task whose success depends on maintaining reliable behaviour across many actions, phases, or dependent subtasks. The term has no universal step-count threshold: it usually signals sequential dependencies, accumulating execution error, delayed outcomes, changing object state, or information that must be remembered beyond the current observation.
Models & learning
Policy
A policy is the decision rule that maps a robot’s current observations or estimated state, and sometimes a task instruction, to an action or probability distribution over actions. It can be hand-designed or learned from demonstrations, rewards or both. In humanoid robotics, its outputs may be joint targets, torques, end-effector changes or higher-level skills.
Data & collection
Trajectory augmentation
Trajectory augmentation is the creation of additional robot trajectories by transforming, replaying, generating, or re-executing existing trajectories under changed conditions while preserving defined task semantics. Useful augmented trajectories are validated for feasibility and outcome rather than accepted solely because their coordinates were transformed.