humanoidsdata.com

Search

Search companies, datasets, articles, and glossary terms for humanoids and embodied AI.

By Lumi · Humanoid robot data · · 8 min read

What Humanoid-GPT’s 2 Billion Motion Frames Actually Teach

Humanoid-GPT shows that scaling a carefully processed motion corpus and a causal Transformer can improve how one Unitree G1 tracks unseen reference motions. It does not show a robot independently choosing a task, understanding a scene, or manipulating an object. The distinction matters: this is a substantial result about general-purpose whole-body tracking, not yet a general-purpose humanoid controller.

The CVPR 2026 paper reports a 92.58% success rate for its largest model on an unseen simulation test split, plus real-robot tracking of held-out dance motions and live motion-capture input. “Success” in that table means staying upright while following a supplied trajectory. It is not household-task completion, and the figures are author-reported rather than independently replicated.

Unitree G1 humanoids tracking varied reference motions with the Humanoid-GPT controller

Humanoid-GPT is evaluated as a whole-body reference-motion tracker on the Unitree G1. Source: the authors’ project page and CC BY 4.0 paper.

Tracking a motion is not choosing a task

In motion tracking, the controller receives a target pose or trajectory and tries to make the robot follow it while remaining dynamically stable. Humanoid-GPT takes the current G1 proprioceptive state and a retargeted reference pose, then predicts per-joint proportional–derivative targets. A lower-level controller turns those targets into actuator torques.

That interface is narrower than an autonomous robot policy. The tracker does not infer that a room needs disinfecting, decide where to dig, or plan which object to grasp. Another system—or a live human motion stream—must supply the reference. The paper’s future-work section explicitly leaves contacts, vision, language, interactive scenarios, and longer-horizon planning for later work.

This is also why “zero-shot” needs a precise object. The authors report tracking motions excluded from training without motion-specific fine-tuning. They do not report zero-shot transfer to a new robot body: training data, reference targets, actions, and evaluation are all built around the 29-degree-of-freedom Unitree G1.

A G1 follows live motion-capture references in a home setting. The video demonstrates tracking and balance, not autonomous task selection. Source: the Humanoid-GPT project.

Two billion frames are not two billion robot interactions

The number describes processed motion frames, not two billion independently collected robot episodes. The authors aggregate AMASS, LAFAN1, Motion-X++, PHUMA, MotionMillion, and in-house recordings, then map human motion into the G1’s joint space. They filter sequences such as sitting, swimming, and stair climbing that do not fit the paper’s plain-scene setting, and use time-warping to expand temporal variation to roughly five times the original corpus.

That pipeline changes what each training unit means. A retargeted frame contains a desired robot configuration derived from human motion. It does not contain a camera observation, language instruction, object state, contact force, or task outcome unless those signals are added separately. Our guide to human motion as robot training data explains why the capture, retargeted reference, controller action, and measured execution should remain distinct in a dataset record.

The source mix matters as much as the count. Common movements can dominate a large corpus while rare balance transitions disappear in the long tail. Humanoid-GPT introduces Harmonic Motion Embedding to group sequences by periodic joint features, producing roughly 300 clusters with about 1,000–2,000 sequences each. The clusters organize specialist training and help balance the distribution; they are a proposed representation from this paper, not a standardized measure of motion diversity.

Humanoid-GPT pipeline from curated and retargeted motion data through reinforcement-learning experts to a distilled causal Transformer

The pipeline converts human motion into G1 references, trains clustered reinforcement-learning experts, and distils their actions into one causal Transformer. Source: the project page and paper.

The corpus trains experts before it trains the Transformer

Humanoid-GPT is not trained by directly copying all two billion target poses. The paper describes four linked stages:

  1. Human motions are cleaned, augmented, and retargeted into G1 reference trajectories.
  2. The embedding clusters related motion sequences.
  3. About 384 reinforcement-learning experts learn to track individual clusters in simulation using rewards for keypoint position, rotation, velocity, stability, and smoothness.
  4. One causal Transformer learns the experts’ actions through policy distillation with DAgger-style data collection.

The fourth stage is especially important. During distillation, the student executes actions and visits its own states; the appropriate expert supplies a target action for those states. This is closer to the correction loop described by DAgger than to one-pass behaviour cloning on a fixed demonstration file. It exposes the student to some of the mistakes created by its own policy.

The appendix reports approximately 15,000 GPU hours: about 12,000 RTX 4090 hours for specialist training and 3,000 H100 hours for distillation. Only the distilled policy is needed at deployment, but the training bill shows why the final checkpoint should not be mistaken for the whole reproducible pipeline.

Data passed between stagesWhat it teachesWhat it does not establish
Human motion sourcesA broad prior over body movementFeasibility on the G1
Retargeted G1 referencesDesired robot-specific poses and trajectoriesStable execution under dynamics
Expert simulation rolloutsCorrective actions for clustered motion regimesReal-world robustness by itself
DAgger teacher labelsHow specialists would act in states visited by the studentAutonomous goals, perception, or planning
Hardware execution logsWhether the G1 followed a supplied referenceCompletion of an independently selected task

For a training-data team, these layers should have separate provenance. Calling all of them “motion data” hides which parts were recorded, generated, simulated, or measured on hardware.

What the scaling result measures

The cleanest evidence is the paper’s controlled simulation comparison on an unseen AMASS test split. With the architecture held to the small model, increasing training tokens from 2 million to 20 million raised reported tracking success from 83.26% to 86.02%. The 22.1-million-parameter base model reached 88.27% at 200 million tokens and 90.43% at two billion. At the same two-billion-token scale, the 80.4-million-parameter model reached 92.58%.

Humanoid-GPT data scaling curves for zero-shot motion-tracking success and error metrics

The authors report monotonic gains across their tested data sizes, with smaller marginal gains from 200 million to two billion tokens. Source: the Humanoid-GPT paper.

Those results support a data-and-model scaling relationship within this tracker, corpus, robot, simulator, and evaluation. They do not say that multiplying frames by ten multiplies task success, nor that a model will extrapolate indefinitely. The authors themselves note smaller marginal gains between 200 million and two billion tokens, suggesting that the base model was becoming capacity-limited.

This differs from Figure’s reported Helix 2.5 robot-data scaling result. Figure varied human video pretraining and measured held-out action-prediction loss before testing three household tasks. Humanoid-GPT varies retargeted motion frames and model capacity, then measures reference-tracking stability and error. Both are scaling studies, but they scale different data and answer different questions.

The real-robot evidence is useful but bounded

The paper evaluates four held-out dance sequences on a physical G1 and reports joint-position and joint-velocity errors. It also shows live motion-capture retargeting, plus qualitative examples of digging, football, balance recovery, and other movements. These demonstrations show that the simulated tracker can transfer to hardware and follow varied references in real time.

The digging sequence is evidence of reference-motion execution. It does not demonstrate perception-driven digging or measure a completed excavation task. Source: the Humanoid-GPT project.

There are still important limits:

  • The quantitative hardware table covers four dance motions on one robot embodiment. It does not report a broad physical success rate across the full test corpus.
  • Joint-sensor agreement measures how closely the robot followed commanded motion, but it is not an external motion-capture accuracy check and does not measure object or environmental outcomes.
  • The paper filters explicit object-interaction sequences from the training corpus, so its evidence does not establish contact-rich manipulation.
  • The reported latency comes from an optimized deployment using an RTX 4090, an Intel i9-14900KF, ONNX, and TensorRT. It should not be generalized to every onboard compute stack.
  • The authors’ repository currently provides inference and deployment code plus a checkpoint under Apache 2.0, while listing training code and training data as future releases. Reproducing the reported training run is therefore not yet possible from the public repository alone.

What a reusable motion corpus should preserve

Humanoid-GPT makes a strong case that scale becomes useful only after curation, embodiment mapping, balanced sampling, and corrective supervision. A buyer or builder should therefore ask for more than a frame count:

  • Count unique source clips, duration, performers, motion families, and augmentation separately from retargeted frames.
  • Preserve source-dataset and in-house recording provenance, licenses, exclusions, and deduplication rules.
  • Version the retargeter, robot model, joint ordering, coordinate frames, contact assumptions, and filters that produced each reference trajectory.
  • Keep the expert identifier, simulation configuration, student-visited state, teacher action, and rollout seed for distilled supervision.
  • Define each held-out axis—source clip, motion family, performer, environment, or embodiment—and report falls as well as average tracking error.
  • Validate with external pose or task measurements when the claim extends beyond joint-level reference following.

The practical lesson is not simply that two billion frames beat two million. Humanoid-GPT works by turning a large, uneven human-motion collection into balanced robot-specific references, physically trained specialists, and corrective labels for one deployable sequence model. That is a credible route toward a broadly reusable motion primitive. Vision, contact-rich interaction, task planning, new embodiments, and a fully reproducible training release remain separate steps.