Models & learning
Policy distillation
Policy distillation transfers behaviour from one or more teacher policies into a student policy by training the student to match teacher outputs on sampled states. In robotics it can consolidate specialist controllers into one deployable policy, but the student remains limited by the states it visits and the quality and coverage of its teachers.
Also known as: policy knowledge distillation, teacher-student policy distillation
Updated
Teacher behaviour becomes training supervision
A teacher policy maps observations or states to actions. Policy distillation records or queries those outputs and trains a student to reproduce them, often with a supervised loss. The original Policy Distillation paper showed how this teacher–student process could compress policies and combine several task-specific teachers into one model.
The student does not inherit a teacher’s behaviour automatically. Its architecture, observations, action representation, loss, and training-state distribution determine what can be transferred. A small student may lose rare behaviours, and conflicting teachers can provide inconsistent targets.
Why the student’s visited states matter
Training only on states generated by an expert can leave the student without labels for its own mistakes. DAgger addresses that distribution shift by running the learner, asking an expert how it should act in the states the learner visits, adding those labels, and retraining.
Policy distillation can use the same loop. The resulting dataset should preserve the learner state, executed action, teacher identity, teacher target, and iteration. Without that provenance, it is difficult to tell whether the student learned recovery behaviour or only copied ideal trajectories.
Combining motion specialists
The Humanoid-GPT paper trains reinforcement-learning experts on clusters of retargeted motion, then distils their outputs into one causal Transformer with DAgger-style supervision. The specialist library is needed during training but can be discarded at deployment.
That makes deployment simpler without making the teachers irrelevant. The final policy’s capabilities still reflect which motion clusters had competent experts, which student-visited states were labelled, and how conflicts or gaps across the specialist set were handled.
Sources
Related terms
Models & learning
Behaviour cloning
Behaviour cloning is a form of imitation learning that fits a policy to expert observation–action pairs as a supervised prediction problem. For humanoid robots, the training examples typically align camera or proprioceptive observations with commands recorded during demonstrations, so the learned policy can reproduce similar behaviour without an explicit reward model.
Models & learning
Imitation learning
Imitation learning is a family of methods that learns a policy from examples of expert behaviour rather than specifying every control rule by hand. In robotics, demonstrations pair observations or states with actions, trajectories or inferred objectives. Behaviour cloning is one imitation-learning method; interactive and inverse approaches address different supervision and distribution-shift problems.
Models & learning
Policy
A policy is the decision rule that maps a robot’s current observations or estimated state, and sometimes a task instruction, to an action or probability distribution over actions. It can be hand-designed or learned from demonstrations, rewards or both. In humanoid robotics, its outputs may be joint targets, torques, end-effector changes or higher-level skills.
Models & learning
Reinforcement learning
Reinforcement learning is a method in which an agent learns a policy by interacting with an environment and optimising cumulative reward. In humanoid robotics, actions change the robot and world, while observations, rewards and episode endings provide experience for improving balance, locomotion or manipulation behaviour.