humanoidsdata.com

Search

Search companies, datasets, articles, and glossary terms for humanoids and embodied AI.

← All glossary terms

Models & learning

Robot generalization

Robot generalization is a policy's ability to perform a learned behavior when test conditions differ from training, such as with new objects, environments, tasks, or robot embodiments. A useful claim specifies which conditions were held out.

Also known as: generalization, out-of-distribution generalization

Updated

What can change at test time

A robot policy generalizes when it continues to perform after something about the test setting changes. The difference may be a new object, room layout, instruction, robot body, camera viewpoint, or task sequence. These are separate forms of transfer, so a result should name the conditions that were held out.

The RT-1 study examined generalization in real-world robot tasks as the training data and its diversity changed. The Open X-Embodiment work explored transfer across data from multiple robot embodiments. These examples test different kinds of variation; success in one setting does not establish generalization to all of them.

What zero-shot means in a robot test

“Zero-shot” usually means the model receives no task-specific adaptation using the target test conditions. It does not necessarily mean the model has never learned the task or the relevant skills. Figure's Helix 2.5 report, for example, describes unseen homes and manipulated objects, while the task behaviors were specified using data collected elsewhere.

For a meaningful comparison, a report should say whether the task, objects, environments, embodiment, or some combination was held out. It should also state if the model was fine-tuned, prompted, or selected using information from the evaluation setting.

How to evaluate it

Evaluation splits should hold out entire scenes or objects when those are the intended transfer conditions. Each trial needs a defined start state and full-task success rule. Report partial progress, human interventions, failures, and trial counts alongside a headline rate so readers can tell what the policy completed on its own.

Generalization is also distinct from a scaling law: a model may show a predictable change in an intermediate metric without the same change in task success. See Figure's Helix 2.5 data-scaling results for an example.

Sources