Models & learning
Robot generalization
Robot generalization is a policy's ability to perform a learned behavior when test conditions differ from training, such as with new objects, environments, tasks, or robot embodiments. A useful claim specifies which conditions were held out.
Also known as: generalization, out-of-distribution generalization
Updated
What can change at test time
A robot policy generalizes when it continues to perform after something about the test setting changes. The difference may be a new object, room layout, instruction, robot body, camera viewpoint, or task sequence. These are separate forms of transfer, so a result should name the conditions that were held out.
The RT-1 study examined generalization in real-world robot tasks as the training data and its diversity changed. The Open X-Embodiment work explored transfer across data from multiple robot embodiments. These examples test different kinds of variation; success in one setting does not establish generalization to all of them.
What zero-shot means in a robot test
“Zero-shot” usually means the model receives no task-specific adaptation using the target test conditions. It does not necessarily mean the model has never learned the task or the relevant skills. Figure's Helix 2.5 report, for example, describes unseen homes and manipulated objects, while the task behaviors were specified using data collected elsewhere.
For a meaningful comparison, a report should say whether the task, objects, environments, embodiment, or some combination was held out. It should also state if the model was fine-tuned, prompted, or selected using information from the evaluation setting.
How to evaluate it
Evaluation splits should hold out entire scenes or objects when those are the intended transfer conditions. Each trial needs a defined start state and full-task success rule. Report partial progress, human interventions, failures, and trial counts alongside a headline rate so readers can tell what the policy completed on its own.
Generalization is also distinct from a scaling law: a model may show a predictable change in an intermediate metric without the same change in task success. See Figure's Helix 2.5 data-scaling results for an example.
Sources
Related terms
Models & learning
Policy
A policy is the decision rule that maps a robot’s current observations or estimated state, and sometimes a task instruction, to an action or probability distribution over actions. It can be hand-designed or learned from demonstrations, rewards or both. In humanoid robotics, its outputs may be joint targets, torques, end-effector changes or higher-level skills.
Models & learning
Robot foundation model
A robot foundation model is a broadly pretrained model intended to provide a reusable starting point for multiple robot tasks, environments or embodiments. It learns from diverse robotics and sometimes web or human data, then acts directly or is adapted with target-domain data. The term describes a training and reuse strategy, not one fixed architecture.
Models & learning
Scaling law
A machine-learning scaling law is an empirical relationship between training scale, such as dataset size, model size, or compute, and a measured outcome, commonly model loss. It can guide estimates within the tested setup but does not guarantee similar gains for a different metric or deployment setting.