humanoidsdata.com

Search

Search companies, datasets, articles, and glossary terms for humanoids and embodied AI.

By Lumi · Humanoid robot data · · 7 min read

In-Context Learning for Robots: What Skild AI’s S1 Does With One Video

In its S1 research post, Skild AI describes a pretrained robot policy that uses one video demonstration as a prompt: show the task, and the same model weights attempt to carry it out without task-specific fine-tuning. That is a robotics example of in-context learning, not a robot learning manipulation from scratch from one clip.

The post describes tests on new, long-horizon manipulation tasks, including plant potting, pancake cooking, pour-over coffee, and kit assembly. Its reported results are promising, but come from internal benchmarks and include human recovery interventions during scoring. The headline 66% figure is an average per-step result, not the share of complete tasks finished autonomously.

What the video does for S1

In the classic language-model formulation, examples in a prompt can guide a model's answer without changing its weights. S1 applies that idea to robot control. The video supplies a concrete task example; the pretrained model uses it to infer the goal, objects, sequence, and progress, then predicts actions for the robot. The task video is context at inference time, not a fine-tuning dataset.

That distinction changes how to read “one video.” S1 still depends on broad pretraining. Skild describes training on episodic data where tasks are specified through demonstrations, and says its model learns to map demonstrations to actions across different scenes, viewpoints, and embodiments. The single clip specifies a new task for a model that already has learned behavior.

A person records a task demonstration video to use as a prompt for a robot policy

Skild AI shows a person recording a demonstration to use as S1's prompt. Source: Skild AI's S1 research post.

“Unseen” has more than one meaning

Skild's S1 research post evaluates in-context learning along two axes: whether the task appeared in pretraining and how long the task takes. It says its four long-horizon demonstrations were absent from the pretraining task set. The examples—plant potting, making pancakes, pour-over coffee, and assembling a kit—take up to ten minutes and span dozens of steps.

That is different from proving the robot has never seen any part of the behavior before. The research describes the model composing manipulation primitives learned during pretraining into new sequences. A new task may therefore be unfamiliar as a whole while still using familiar grasps, object interactions, or motions. For buyers, “unseen” should be broken down into the held-out task, objects, environment, and robot embodiment.

The video-prompt approach is also distinct from the single-video real-to-sim workflow. SimFoundry uses a video to reconstruct a scene for simulation and further training or evaluation. S1 uses a demonstration as context for the robot policy at execution time. Both start with video, but the video does a different job.

What the published scores show

In its S1 research post, Skild reports a comparison between S1's demonstration-conditioned policy and a language-conditioned vision-language-action model. The company says both used the same filtered pretraining data, architecture apart from the prompt embedding, and compute, with training scales from 1,000 to 100,000 hours. It evaluated two internal benchmark suites: one with tasks from the pretraining distribution and one with unseen tasks.

On the unseen suite at the 100,000-hour scale, Skild reports a 66% average cumulative per-step success rate for S1 and 9% for the language-conditioned baseline. Those tasks ran for four to eight minutes. The score is not the percentage of complete tasks finished without help: the researchers used human interventions to recover from errors so they could score all steps. They say this was mainly necessary for the baseline, which otherwise did not complete a full long unseen task in their test.

Skild also estimates that one in-context demonstration performed like roughly 380 task-specific post-training demonstrations. That figure came from interpolating between measured results; it is specific to the study's tasks and pretraining setup. The same post reports that 2,000 post-training demonstrations eventually reached an 86% success rate, above S1's single-demonstration result. The comparison suggests a potentially large reduction in task-specific collection, not a universal exchange rate between one video and 380 robot episodes.

Skild AI's published pour-over coffee example. The company presents this as an S1 task demonstration; it is not an independent evaluation. Source: Skild AI's S1 research post.

These are company results rather than an independent replication. Skild's post describes internal benchmark suites and does not link to downloadable S1 weights or a public copy of the benchmark. A buyer should treat the reported numbers as a reason to run a matched acceptance test, not as a deployment guarantee.

The data supply does not disappear

If a model can use a short task video, it may reduce the number of robot episodes needed for each new task. It does not remove the need for a substantial training set. In its data-engine discussion, Skild says robot teleoperation is close to robot hardware but difficult to scale, while egocentric video scales more easily but has a larger gap from robot behavior. Its approach is to combine sources rather than rely on one.

NVIDIA's September overview of S1 describes Skild's broader training pipeline as using simulation, human video, teleoperation, and deployment data where customer agreements permit. That is a vendor account of one company's pipeline, not evidence that all models can use the same mix. It does underline the distinction between a single video used at deployment and the data that taught the base model how to interpret demonstrations.

For data teams, the likely shift is from counting only task-specific episodes to asking how broad pretraining, prompt quality, action mapping, and evaluation work together. The robot training data modalities guide explains what video, state, action, contact, and provenance contribute. The guides to robot teleoperation systems, simulation environments, and human motion data cover three of the source types now discussed in the S1 pipeline.

Questions to ask before buying a one-video workflow

  • What exactly was held out: the task sequence, objects, room, robot, or all four?
  • Does “success” mean a completed task, a per-step score, or a score that includes human recovery?
  • Which robot, hand, controller, cameras, and action representation were used in the published test?
  • What does the prompt video need to show, and can the buyer inspect the video, robot log, and outcome labels for a test episode?
  • Are customer demonstration videos stored, reused for model training, or shared with other customers? What rights, retention, and deletion rules apply?
  • Can the vendor run a held-out test on the buyer's robot, with autonomous task completion and intervention counts reported separately?

These questions connect video prompting to the same evidence buyers already need when they evaluate humanoid robot training data or estimate how many demonstrations a robot policy needs. A prompt video may be cheap to record, but its value depends on the model's prior training, the robot's embodiment, and a test that shows what happens when the sequence breaks.

The practical implication

S1 makes a useful case for treating a task video as an input to a pretrained robot brain, rather than only as raw material for a new training run. If the reported results transfer, teams may spend less time collecting a separate set of robot episodes for every task. The cost then moves toward broad pretraining, prompt design, safety checks, and proving performance on the target robot.

For now, the reported 66% is a per-step score from an internal test with human recovery, and the 380-example comparison is an interpolated result on that study's tasks. Those claims are worth testing. They do not yet show that one video can reliably teach an arbitrary humanoid a new job.