Models & learning
Offline reinforcement learning
Offline reinforcement learning is reinforcement learning in which a policy is learned from a fixed dataset of previously collected interactions without collecting additional environment interactions during training. The central difficulty is evaluating and improving actions that may be poorly represented or absent in that dataset.
Also known as: offline RL, batch reinforcement learning, batch RL
Updated
Learning is limited to a fixed record
Levine and colleagues define offline reinforcement learning around learning from a previously collected dataset without further interaction. This is attractive for humanoids because online exploration can be slow, costly and unsafe.
The dataset may mix demonstrations, autonomous trials, failures and several behaviour policies. Unlike behaviour cloning, offline RL uses rewards and transition structure to optimise beyond directly copying the recorded actions.
Distribution shift is the central risk
A learned policy can propose actions outside the dataset's support. A value function trained only on recorded transitions may then assign unrealistic values to unfamiliar actions, and policy optimisation can exploit those errors. Many offline-RL methods constrain the learned policy, penalise uncertainty or use conservative value estimates.
These methods do not create evidence for unseen behaviour. Narrow, repetitive data still limits what can be learned, especially for contacts, falls and recovery states that are rare in successful demonstrations.
Dataset documentation determines reuse
An offline-RL release should preserve observations, actions, rewards, next observations, episode boundaries and termination causes. It should identify the collection policies, exploration process, task versions and whether timeouts were marked as terminal failures.
Report coverage as well as volume: tasks, objects, environments, state regions, action ranges and success/failure balance. Final evaluation should occur in a separate simulator or on appropriately supervised hardware. High estimated return on the fixed dataset is not evidence of safe real-robot performance.
Sources
Related terms
Models & learning
Reinforcement learning
Reinforcement learning is a method in which an agent learns a policy by interacting with an environment and optimising cumulative reward. In humanoid robotics, actions change the robot and world, while observations, rewards and episode endings provide experience for improving balance, locomotion or manipulation behaviour.
Models & learning
Policy
A policy is the decision rule that maps a robot’s current observations or estimated state, and sometimes a task instruction, to an action or probability distribution over actions. It can be hand-designed or learned from demonstrations, rewards or both. In humanoid robotics, its outputs may be joint targets, torques, end-effector changes or higher-level skills.
Models & learning
Reward function
A reward function maps a transition, state, action, outcome, or related task information to a scalar signal used to define what a reinforcement-learning agent should optimise. It encodes an objective for learning; it is not a guarantee that the resulting robot behaviour matches the designer's broader intent.
Models & learning
Behaviour cloning
Behaviour cloning is a form of imitation learning that fits a policy to expert observation–action pairs as a supervised prediction problem. For humanoid robots, the training examples typically align camera or proprioceptive observations with commands recorded during demonstrations, so the learned policy can reproduce similar behaviour without an explicit reward model.
Data & collection
Demonstration
A demonstration is a recorded example of how an intended task or behaviour is performed, usually represented as a time-aligned sequence of observations, states and actions. For humanoid robot learning, demonstrations may come from teleoperation, kinaesthetic guidance, motion capture or autonomous experts and provide targets for imitation.
Data & collection
Robot training data
Robot training data is recorded experience used to train, fine-tune, or adapt models for robot perception, prediction, planning, or control. It can include sensor observations, robot state, actions, task instructions, rewards or outcomes, demonstrations, failures, and embodiment metadata. Not every dataset contains every field, but their timing and physical meaning must be clear.