humanoidsdata.com

Search

Search companies, datasets, articles, and glossary terms for humanoids and embodied AI.

By Lumi · Humanoid robot data · · 6 min read

What Figure's Helix 2.5 Shows About Robot Data Scaling

Figure's 17 September 2026 Helix 2.5 announcement makes two related but different claims about learning from human data. Figure reports that adding more Index pretraining data lowered held-out robot-action prediction loss. In a separate test across three tasks and 30 unseen Bay Area homes, the company reports 56% full-task success for an Index-pretrained policy and 9% for one trained from scratch.

The first result measures a prediction loss as a dataset grows. The second measures whether a robot finishes particular household tasks in new environments. Both are relevant to robot generalization, but only the first is the company's scaling law. Figure's results are company-reported, so they are best read as a detailed claim to test rather than an independent replication.

What the scaling study measures

Figure says it trained four models on nested subsets of Index, spanning an eightfold increase in pretraining data. Model size and downstream task training were held fixed. After fine-tuning each model on the same task data, Figure measured held-out robot-action prediction loss. The loss fell with each doubling of Index data, and the company says it could forecast the largest run's test loss from the smaller runs before training it.

The reported forecasting error was 0.54% of the variation across the full eightfold data range. That figure describes the fit to the tested loss curve. It is not a 0.54% error on household task success, and the study does not show what happens outside its tested range or with a different model size.

The Index launch post describes the pretraining source as videos that people record while doing tasks at home and at work. Figure reports that each 1,000 collected hours contains 373 unique tasks, 1,146 unique manipulated objects, and 116 unique environments. The company also describes filtering, fraud review, deduplication, rebalancing, and captioning its uploads. Those are useful details about how Index is assembled, but they are company-reported counts rather than an independent audit of the data's usefulness for a particular robot.

Figure's chart of held-out robot-action prediction loss across increasing amounts of Index pretraining data

Figure's chart for its human-to-robot transfer scaling study. Its reported 0.54% forecasting error is relative to loss variation across the tested data range. Source: Figure's Helix 2.5 announcement.

The 30-home test measures another outcome

Figure evaluated Helix 2.5 on tidying living rooms, folding towels, and making beds in 30 Bay Area homes. The homes and manipulated objects were held out: the company says it collected no data in the evaluation homes, and task-specification data did not include the evaluation objects. A single fixed checkpoint for each task ran across all 30 homes, with no fine-tuning or checkpoint selection using those evaluation rollouts.

Here, “zero-shot” describes the homes and objects, not an entirely new task with no task-specific robot training. Figure says the behaviors were specified with fine-tuning data collected elsewhere. That distinction matters: a robot may face a new layout and new objects while still relying on task-specific examples gathered in other settings.

To isolate the contribution of Index pretraining, Figure compared policies trained on identical task-specific data. One started from random weights; the other started from Helix 2.5's Index-pretrained weights. The announcement says architecture, optimization, hyperparameters, downstream data, and evaluation were held fixed. It reports 9% success for the policy trained from scratch and 56% for the Index-pretrained policy. Success meant finishing the entire task, such as placing all scattered toys in a basket, rather than receiving partial credit.

Figure's published examples of Helix 2.5 on its three whole-body tasks. The clip illustrates the company's report; it is not an independent evaluation. Source: Figure's Helix 2.5 announcement.

These results should not be collapsed into one measure. The scaling study tracks held-out action prediction. The home study tracks complete task outcomes under a specific protocol. A better prediction loss may help downstream behavior, but it does not by itself prove that full-task success will rise at the same rate.

The novelty claim has a specific boundary

Figure describes its result as the first human-to-robot transfer scaling law measured on a humanoid. Read narrowly, that claim is about scaling human-data pretraining for humanoid transfer. It is not a claim that robotics has never studied scaling. The earlier RT-1 paper tested how robot performance varied with dataset size, model size, and data diversity.

The difference is the training source and what is being scaled. RT-1 studied robot data and robot policies. Figure says Helix 2.5 was pretrained on human behavior data from Index, then fine-tuned on robot task data. Its announcement frames the contribution as a human-to-humanoid transfer curve. That is a narrower and more useful comparison than treating this as the first robotics scaling result of any kind.

The data path is also different from Skild's in-context learning result, where a demonstration video conditions a pretrained policy at execution time. Figure uses Index videos during pretraining, then uses robot task data to specify the behaviors. For more on the separate route from human video to robot-specific motion, see how human motion can become robot training data.

What data teams should test next

Figure's experiment points to a useful question for a training-data team: if the model and downstream task data stay fixed, does a larger, more diverse human dataset improve the robot's held-out behavior? A careful test should track both the training metric and the deployed outcome.

  • Keep the model, task-specific data, tuning procedure, and evaluation rules fixed when measuring the effect of one new data source.
  • Hold out whole homes, objects, or work sites, then report what changed between training and evaluation. “Unseen” is not precise unless the held-out factor is named.
  • Report end-to-end completion, partial progress, attempts, failures, human interventions, and the number of trials. A prediction loss and a complete-task success rate answer different questions.
  • Check whether the dataset contains the right signals for the target policy. The Index launch describes human video uploads; that is different from time-aligned robot state, control commands, contact signals, and executed outcomes.
  • Verify collection rights and usage terms, especially when videos include people, homes, or workplaces.

The guides to evaluating humanoid robot training data, robot demonstration counts, and the robotics data pyramid cover the buyer questions behind those checks.

Helix 2.5 is an encouraging company-reported case for broad human pretraining, with two distinct pieces of evidence: lower held-out action-prediction loss as Index grows, and higher task success than a scratch-trained comparison in a held-out-home test. The next useful evidence is a reproducible protocol and matched performance on the target robot. The scaling curve for prediction loss should not be treated as a scaling curve for every task a humanoid might face.