Models & learning
Reward function
A reward function maps a transition, state, action, outcome, or related task information to a scalar signal used to define what a reinforcement-learning agent should optimise. It encodes an objective for learning; it is not a guarantee that the resulting robot behaviour matches the designer's broader intent.
Also known as: reward signal, robot reward function
Updated
Reward defines the optimisation target
In reinforcement learning, an agent seeks to maximise expected cumulative reward rather than each immediate reward independently. Sutton and Barto use reward to formalise the goal while separating it from the agent's internal value estimates.
A robot reward can combine task completion, progress, energy, speed, smoothness, collision penalties or safety constraints. The weighting and time horizon determine the trade-offs the policy sees.
Sparse and shaped rewards make different compromises
A sparse reward may score only successful completion. It is easy to interpret but can provide little learning signal. Reward shaping adds intermediate terms, such as distance to an object, yet can introduce shortcuts: a policy may optimise the proxy without completing the intended task.
A learned reward model estimates reward from data or preferences. A reward function is the broader mapping used by the learning problem and may be hand-coded, learned or hybrid. Neither should be confused with a benchmark metric reported after training.
Reward metadata is part of dataset provenance
Offline datasets often contain stored rewards, but the numbers are uninterpretable without the formula, units, termination rules, discount convention and software version. Clipped or normalised rewards may discard distinctions required for a new objective.
For physical robots, report constraints separately when violations are unacceptable. A large collision penalty still permits collision in the mathematical objective if another return is larger. Evaluation should inspect behaviour and safety outcomes, not only accumulated reward.
Sources
Related terms
Models & learning
Reinforcement learning
Reinforcement learning is a method in which an agent learns a policy by interacting with an environment and optimising cumulative reward. In humanoid robotics, actions change the robot and world, while observations, rewards and episode endings provide experience for improving balance, locomotion or manipulation behaviour.
Models & learning
Reward model
A reward model is a learned function that estimates task progress, success, preference, or another reward signal from robot observations or trajectories. It can replace or supplement a hand-written reward when training or evaluating a policy, but its output is only as reliable as its labels, coverage, and resistance to shortcuts.
Models & learning
Markov decision process
A Markov decision process is a mathematical model of sequential decision-making defined by states, actions, transition probabilities and rewards. After an agent chooses an action, the current state and action determine the distribution of the next state and reward. Reinforcement-learning methods use this structure to compare policies by expected cumulative reward.
Models & learning
Policy
A policy is the decision rule that maps a robot’s current observations or estimated state, and sometimes a task instruction, to an action or probability distribution over actions. It can be hand-designed or learned from demonstrations, rewards or both. In humanoid robotics, its outputs may be joint targets, torques, end-effector changes or higher-level skills.
Models & learning
Long-horizon task
A long-horizon task is a temporally extended robot task whose success depends on maintaining reliable behaviour across many actions, phases, or dependent subtasks. The term has no universal step-count threshold: it usually signals sequential dependencies, accumulating execution error, delayed outcomes, changing object state, or information that must be remembered beyond the current observation.