Models & learning
Scaling law
A machine-learning scaling law is an empirical relationship between training scale, such as dataset size, model size, or compute, and a measured outcome, commonly model loss. It can guide estimates within the tested setup but does not guarantee similar gains for a different metric or deployment setting.
Also known as: machine-learning scaling law, data scaling law, model scaling law
Updated
What the relationship measures
A scaling law summarizes how a chosen model metric changes as one training dimension grows. A study may vary dataset size, model parameters, or training compute. The classic language-model scaling study by Kaplan and colleagues measured how cross-entropy loss changed with model size, dataset size, and compute. Those relationships were measured for particular models and training setups; they are not a guarantee for every system.
Reading a robotics scaling claim
A robot study needs to say which variable changed and which outcome was tracked. Figure's Helix 2.5 announcement says four models were trained on Index subsets spanning an eightfold data increase, with model size and downstream training fixed. The outcome was held-out robot-action prediction loss. Figure reports that the largest run's test loss could be forecast from smaller runs with an error equal to 0.54% of the variation across that tested range.
That result concerns the prediction metric in that setup. It does not show that full-task success improves by the same percentage, or that a curve measured on one model and dataset can be extrapolated to another robot or task. Compare scaling claims with a separate held-out evaluation of the robot policy and the robot generalization conditions that matter in deployment.
Sources
Related terms
Models & learning
Robot foundation model
A robot foundation model is a broadly pretrained model intended to provide a reusable starting point for multiple robot tasks, environments or embodiments. It learns from diverse robotics and sometimes web or human data, then acts directly or is adapted with target-domain data. The term describes a training and reuse strategy, not one fixed architecture.
Models & learning
Robot generalization
Robot generalization is a policy's ability to perform a learned behavior when test conditions differ from training, such as with new objects, environments, tasks, or robot embodiments. A useful claim specifies which conditions were held out.
Data & collection
Robot training data
Robot training data is recorded experience used to train, fine-tune, or adapt models for robot perception, prediction, planning, or control. It can include sensor observations, robot state, actions, task instructions, rewards or outcomes, demonstrations, failures, and embodiment metadata. Not every dataset contains every field, but their timing and physical meaning must be clear.