humanoidsdata.com

Search

Search companies, datasets, articles, and glossary terms for humanoids and embodied AI.

← All glossary terms

Models & learning

Policy distillation

Policy distillation transfers behaviour from one or more teacher policies into a student policy by training the student to match teacher outputs on sampled states. In robotics it can consolidate specialist controllers into one deployable policy, but the student remains limited by the states it visits and the quality and coverage of its teachers.

Also known as: policy knowledge distillation, teacher-student policy distillation

Updated

Teacher behaviour becomes training supervision

A teacher policy maps observations or states to actions. Policy distillation records or queries those outputs and trains a student to reproduce them, often with a supervised loss. The original Policy Distillation paper showed how this teacher–student process could compress policies and combine several task-specific teachers into one model.

The student does not inherit a teacher’s behaviour automatically. Its architecture, observations, action representation, loss, and training-state distribution determine what can be transferred. A small student may lose rare behaviours, and conflicting teachers can provide inconsistent targets.

Why the student’s visited states matter

Training only on states generated by an expert can leave the student without labels for its own mistakes. DAgger addresses that distribution shift by running the learner, asking an expert how it should act in the states the learner visits, adding those labels, and retraining.

Policy distillation can use the same loop. The resulting dataset should preserve the learner state, executed action, teacher identity, teacher target, and iteration. Without that provenance, it is difficult to tell whether the student learned recovery behaviour or only copied ideal trajectories.

Combining motion specialists

The Humanoid-GPT paper trains reinforcement-learning experts on clusters of retargeted motion, then distils their outputs into one causal Transformer with DAgger-style supervision. The specialist library is needed during training but can be discarded at deployment.

That makes deployment simpler without making the teachers irrelevant. The final policy’s capabilities still reflect which motion clusters had competent experts, which student-visited states were labelled, and how conflicts or gaps across the specialist set were handled.

Sources