humanoidsdata.com

Search

Search companies, datasets, articles, and glossary terms for humanoids and embodied AI.

← All glossary terms

Models & learning

Reward function

A reward function maps a transition, state, action, outcome, or related task information to a scalar signal used to define what a reinforcement-learning agent should optimise. It encodes an objective for learning; it is not a guarantee that the resulting robot behaviour matches the designer's broader intent.

Also known as: reward signal, robot reward function

Updated

Reward defines the optimisation target

In reinforcement learning, an agent seeks to maximise expected cumulative reward rather than each immediate reward independently. Sutton and Barto use reward to formalise the goal while separating it from the agent's internal value estimates.

A robot reward can combine task completion, progress, energy, speed, smoothness, collision penalties or safety constraints. The weighting and time horizon determine the trade-offs the policy sees.

Sparse and shaped rewards make different compromises

A sparse reward may score only successful completion. It is easy to interpret but can provide little learning signal. Reward shaping adds intermediate terms, such as distance to an object, yet can introduce shortcuts: a policy may optimise the proxy without completing the intended task.

A learned reward model estimates reward from data or preferences. A reward function is the broader mapping used by the learning problem and may be hand-coded, learned or hybrid. Neither should be confused with a benchmark metric reported after training.

Reward metadata is part of dataset provenance

Offline datasets often contain stored rewards, but the numbers are uninterpretable without the formula, units, termination rules, discount convention and software version. Clipped or normalised rewards may discard distinctions required for a new objective.

For physical robots, report constraints separately when violations are unacceptable. A large collision penalty still permits collision in the mathematical objective if another return is larger. Evaluation should inspect behaviour and safety outcomes, not only accumulated reward.

Sources