humanoidsdata.com

Search

Search companies, datasets, articles, and glossary terms for humanoids and embodied AI.

← All glossary terms

Models & learning

Offline reinforcement learning

Offline reinforcement learning is reinforcement learning in which a policy is learned from a fixed dataset of previously collected interactions without collecting additional environment interactions during training. The central difficulty is evaluating and improving actions that may be poorly represented or absent in that dataset.

Also known as: offline RL, batch reinforcement learning, batch RL

Updated

Learning is limited to a fixed record

Levine and colleagues define offline reinforcement learning around learning from a previously collected dataset without further interaction. This is attractive for humanoids because online exploration can be slow, costly and unsafe.

The dataset may mix demonstrations, autonomous trials, failures and several behaviour policies. Unlike behaviour cloning, offline RL uses rewards and transition structure to optimise beyond directly copying the recorded actions.

Distribution shift is the central risk

A learned policy can propose actions outside the dataset's support. A value function trained only on recorded transitions may then assign unrealistic values to unfamiliar actions, and policy optimisation can exploit those errors. Many offline-RL methods constrain the learned policy, penalise uncertainty or use conservative value estimates.

These methods do not create evidence for unseen behaviour. Narrow, repetitive data still limits what can be learned, especially for contacts, falls and recovery states that are rare in successful demonstrations.

Dataset documentation determines reuse

An offline-RL release should preserve observations, actions, rewards, next observations, episode boundaries and termination causes. It should identify the collection policies, exploration process, task versions and whether timeouts were marked as terminal failures.

Report coverage as well as volume: tasks, objects, environments, state regions, action ranges and success/failure balance. Final evaluation should occur in a separate simulator or on appropriately supervised hardware. High estimated return on the fixed dataset is not evidence of safe real-robot performance.

Sources