humanoidsdata.com

Search

Search companies, datasets, articles, and glossary terms for humanoids and embodied AI.

← All glossary terms

Models & learning

Curriculum learning

Curriculum learning is a training strategy that changes which examples, tasks, or initial states a model encounters over time, often moving from easier conditions towards harder ones or expanding the distribution as competence improves. In robot learning, the curriculum controls exploration and learning signal; a poorly balanced sequence can undertrain later stages or degrade earlier skills.

Also known as: curriculum training, training curriculum

Updated

A curriculum changes the training distribution

A fixed training distribution samples the same mixture throughout learning. A curriculum deliberately changes that mixture. It may introduce harder tasks later, widen a range of initial states, increase the length of an episode, or sample conditions near a goal before moving farther away.

Reverse Curriculum Generation demonstrates the last approach for reinforcement learning: starts are sampled near a known goal and expanded outwards as the policy succeeds. The useful abstraction is not “easy examples first” by itself. The curriculum specifies which experience becomes available, when it appears, and how progress changes the sampling rule.

Robot curricula can hide missing stages

Curricula help when rewards are too sparse for a robot to discover a complete task from scratch. They can also bias training. If early states deliver reward frequently, the policy may spend most of its updates on them while later transitions receive little useful signal. Moving entirely to later states can create a different problem if new updates damage earlier competence.

Humanoid Horizon reports this imbalance in sequential object transport. Its alternative keeps stage streams active in parallel and updates one shared policy from all of them, while rollout-derived terminal states change the start distribution for downstream stages.

Record the curriculum as data provenance

A robot dataset produced during training should preserve the curriculum state alongside each episode: task or stage, difficulty parameters, initial-state source, sampling probability, generator or policy checkpoint, success rule, and the condition that advanced or reset the curriculum.

Without that record, a large trajectory set can conceal that most episodes came from easy starts or that harder stages were sampled only after a policy became specialised. Evaluation should use a separately declared distribution so success is not measured only on states chosen by the training curriculum.

Sources