The Rundown AI homepage
Robotics/News & analysis

Skild’s S1 aims to cut robot training time with a single video

Skild says S1 can learn an extended robot task from one video. Its tests suggest less task data may be needed, though reliable completion remains uncertain.

By Jennifer Mossalgue3 min read
Skild wants to teach robots with a single video — newsletter story image
Image source: Skild AI

Skild AI says its S1 robot foundation model can watch a human demonstrate an unfamiliar job once, then attempt that job on a robot without updating its model weights. The company’s examples include potting a plant and flipping pancakes, with some tasks lasting up to ten minutes across dozens of steps.

The announcement appeared in The Rundown’s September 3 newsletter. Skild dates S1’s introduction to August 18. Its central promise is to reduce the repeated demonstrations and additional training needed to introduce a robot to a new, extended task.

What S1 demonstrated

According to Skild’s S1 report, the model performed plant potting, pancake preparation, coffee preparation and kit assembly tasks that were absent from its pretraining. The weights stayed unchanged across these demonstrations; the new video supplied the task instructions as an input.

In the plant example, recording the human demonstration began at 9:16 p.m., and autonomous execution began at 9:27 p.m. That explains the 11-minute turnaround highlighted in the September 3 newsletter. Setup had started at 8:54 p.m., so the interval describes the period from recording to execution rather than the complete setup process.

Skild also reports a controlled comparison at 100,000 hours of pretraining: demonstration prompting scored 66%, compared with 9% for language prompting. The definition of that score matters. It averages cumulative success across task steps and includes human recovery interventions. It therefore leaves open how often a robot could complete an entire job without assistance.

The company estimates that matching S1’s single-video performance through additional training would take about 380 teleoperated demonstrations, equivalent to 50–100 hours. That crossover was interpolated between measured results. With 2,000 demonstrations, the additional-training approach reached 86%, exceeding the single-video result on Skild’s metric.

Why it matters

Teaching a robot a new job that takes many steps can still require tens to hundreds of hours of task-specific data. Skild’s comparison puts a concrete scale on that burden: in its experiment, one video prompt delivered performance that the company estimates would otherwise require roughly 380 teleoperated demonstrations. The potential gain is a shorter path from identifying a new task to testing whether a robot can perform it. The hours estimate applies to Skild’s experiment; requirements will depend on the task and training approach.

For teams introducing robots to new workflows, that could change the economics of the first trial. Collecting repeated demonstrations takes operator time before the team can assess the resulting behavior. If a single recorded example provides a useful starting point, teams could spend less time assembling an initial task dataset and begin evaluating the robot sooner. Skild’s results support that possibility, while leaving actual deployment savings unmeasured.

The substantial training investment comes earlier. Skild’s reported comparison involved 100,000 hours of pretraining, and a fresh demonstration still enters the system as a prompt. The proposed efficiency comes from applying that existing foundation to another job with little additional task data. How broadly that carries across unfamiliar environments and hardware remains an open question.

Generalist provides another example of this direction. In its August 19 GEN-1.5 announcement, it reported 59% average success across ten simple, short tasks after receiving 3–12 seconds of one demonstration, without weight updates. Its standard “physical prompts” include sensor data and action trajectories, with separate examples of transfer from bare-hand human demonstrations. Those results offer a different test of rapid adaptation from Skild’s extended tasks.

Both companies’ reports also suggest a role for further training. Generalist reported 83% success after ten gradient steps with five minutes of task data, while Skild’s larger demonstration collection exceeded its single-video score. A practical approach could be to demonstrate a task, assess performance, then collect targeted data where reliability falls short.

That makes full-job testing essential. Skild’s step-based score, which includes recovery assistance, leaves the frequency of successful unassisted completion unresolved. Before assigning a workflow to S1, a team would need to measure complete runs, interventions and recovery behavior under the conditions where the robot will work.

Sources & further reading

This story builds on reporting from The Rundown newsletter on September 3, 2026.