Imagine a factory suddenly telling you:

"Starting tomorrow, this robot no longer assembles part A."

"It will assemble a different part instead."

"Also, the sequence has changed."

For a person,

this might mean a supervisor stands beside you and demonstrates once.

You watch, understand,

and then begin to try.

But for many robots,

things aren’t that simple.

The biggest challenge for robots isn’t that they can’t move

Modern robotic arms

can actually already:

pick things up.

place things.

transport.

stack.

assemble.

The problem is:

once the task changes, they often need to relearn.

When introducing the S1 robot, Skild AI stated directly that

traditional robots require collecting several hours of teleoperation demonstration data

and performing task-specific fine-tuning

to deploy a new task.

For complex, unfamiliar tasks involving many steps, even more demonstrations may be needed to improve reliability.

This might be feasible in labs,

but real factories never stay the same.

Factories often face "yet another change"

A supplier changes a part.

Workstations are rearranged.

Product versions update.

Assembly sequences alter.

Packaging methods differ.

Whenever process changes,

previously calibrated robots

might need reconfiguration again.

Skild shared from real deployments:

If every change means recollecting data and post-training,

then for every day a robot works,

you must constantly retrain it.

This is difficult to scale.

So Skild’s challenge is not

"how to make robot hands more dexterous",

but rather:

"how to get robots to quickly learn something new."

Skild’s solution: No retraining, just show it once

The newest robot foundation model from Skild is called

S1.

Its biggest difference

is adopting the AI concept we are familiar with:

In-context Learning.

We now use ChatGPT,

and if we want it to write something in a certain format,

we don’t need to retrain it.

We just give it an example prompt

and say: "Do it like this."

S1 wants to bring the same concept into the real world.

But instead of text prompts,

the prompt it sees is:

A video.

For example: Teaching it to plant a pot

You prepare:

a pot,

some soil,

a plant,

and a watering can.

Then you actually do the task once:

put soil in,

dig a hole,

place the plant,

fill with soil,

and water it.

A camera records the entire process.

In the past,

this video might just be used as training data,

which requires sorting, training a model, testing, and revising.

But S1 works differently.

This video itself is the prompt.

It directly enters the model’s context.

The model does not modify its original weights,

nor does it fine-tune for the new task.

Is it just copying the video’s movements?

No, and this is very important.

If it just replays the exact movements in the video, like:

move hand left 10 cm,

then down 5 cm,

and repeats exactly,

it wouldn’t be that valuable.

The real world is never exactly the same.

The location of the pot might change.

The cup might be somewhere else.

The tools might not be identical.

S1’s job is to understand from the video:

what goal the human is achieving,

what role each item plays,

the relationship between steps,

which step has been completed,

and what comes next.

It then converts this information

into actions for the robot in front of it.

Skild says that the same S1 model weights

can handle both seen and unseen tasks, without needing fine-tuning for each new job.

The most impressive number: 11 minutes

Skild tested S1 planting a pot.

At 9:16 pm, they began recording.

At 9:22 pm, they finished recording a human demonstration video.

By 9:27 pm,

S1 was already trying the task on actual robot hardware.

That is:

only about 11 minutes from demonstration recording to autonomous execution.

This is where the real significance lies.

Not that the robot has "finally learned to plant a pot."

That task itself isn’t critical.

What matters is:

the time needed to teach a new job may drastically shrink.

And these aren’t just quick three-second moves

Many robot demos look impressive.

Picking up a cup.

Placing a part.

Opening a drawer.

These videos may last just a few seconds.

The real challenge is:

linking many actions together.

The longer the process,

the higher the chance of errors.

If something isn’t placed correctly early on,

everything after might go wrong.

Skild’s S1 demos show

tasks lasting up to about 10 minutes,

such as:

planting a pot,

making pancakes,

brewing coffee manually,

assembling kits.

These involve dozens of operational steps,

and all were entirely new tasks not seen during pre-training.

Flipping a pancake is a great example

If I say,

"Flip the pancake,"

a human generally understands what it means.

But for a robot,

language alone doesn’t explain:

where to insert the spatula,

the angle,

wrist rotation,

how much force to apply,

how to catch the pancake mid-air,

and follow-through.

Some physical actions are very hard to describe well with words alone.

That’s why Skild believes,

for complex physical tasks,

directly "showing" the robot may be more effective than "telling" it verbally.

But it’s not yet "watch once and master"

This must be made clear.

Skild ran a set of unfamiliar, multi-step task tests.

With pre-training data scaling to 100,000 hours,

the video demonstration in-context learning model

achieved an average per-step success rate of about:

66%.

By comparison, the language-prompted VLA policy scored around:

9%.

66% is much higher than 9%,

but 66% is not 100%.

In other words:

we can’t think of S1 as

"just film once, and the factory never needs an engineer again."

There are still many challenges remaining

before truly large-scale, long-term, high-reliability production can be achieved.

Skild even allows for human intervention

When testing these unfamiliar long-process tasks,

to allow full completion and accurate measurement of each step,

human intervention was permitted after failures

to resume the process and continue testing.

This means the 66% represents

an average cumulative per-step success rate, not

"the robot can complete 66% of tasks fully unsupervised."

These two concepts shouldn’t be confused.

So why does it still matter?

The real comparison isn’t between

66% and a perfect 100%,

but

how much the cost of teaching new tasks can be reduced.

Skild tested this:

Reaching the 66% performance of S1 with traditional post-training methods

required roughly

380 fine-tuning demonstrations.

These 4–10 minute long demos accumulated approximately

50–100 hours of teleoperation.

S1 instead only used:

one video demonstration.

That’s the key difference.

Robots are entering their own "Prompt Era"

One of generative AI’s biggest changes

is that ordinary people don’t need to retrain a model.

You just tell it:

what you want,

provide some context,

and an example prompt,

then it starts doing it.

Now, physical AI is trying to follow the same path.

In the future, factory supervisors and technicians

may not need to create a new robot training project

every time the process changes.

Some tasks might only require:

doing the new process once,

recording it,

giving the video to the robot,

letting it try.

Then humans can review and correct errors.

This could truly accelerate robots entering factories

Factories have never lacked machines that repeat a fixed task.

Traditional robotic arms

were already great at fixed tasks.

The true challenge is:

when things change, the robot can’t adapt.

Parts move.

Robots don’t adapt.

Processes change.

Robots need reprogramming.

New products arrive.

Robots require retraining.

If in the future robots can quickly adapt from a single demonstration,

the real revolution won’t be any particular robotic arm,

but rather:

how factories train robots.

This might also make "human demonstration skills" more valuable

We used to ask:

Will humans need to learn prompt writing in the future?

If physical AI keeps moving this way,

factories might add a new kind of prompt:

"Show it by doing it yourself."

The clearer the demonstration and the clearer the goals,

the better chance the robot has

to understand what you want to achieve.

At that time,

the key skill might not only be operating robots,

but also:

breaking down tasks clearly,

knowing the correct sequence,

recognizing the most critical steps,

and demonstrating an efficient workflow.

These have always been the expertise of skilled workers.

But in the future,

they might not just teach newcomers,

but also:

teach robots.

So Skild S1’s real significance isn’t mastering pancakes

Pancakes are eye-catching.

Planting pots is visually appealing.

Pour-over coffee makes a great demo.

But the real takeaway

is this:

The gateway for robots learning new jobs is shifting from "retraining models" to "providing context."

S1 currently achieves about 66% success per step on unfamiliar tasks.

It’s not yet fully reliable.

But as this approach improves,

we may soon see a rare scene in factories:

A skilled worker stands at a workstation,

performs a task once,

a robot watches the demonstration,

and then the robot starts doing the job.

At that point,

the way we teach AI

will truly move from typing prompts on a keyboard

to interacting directly in the real world.

Who do you think will be most valuable in the future?

Not "who works faster than robots,"

but "who knows how to teach robots best."

Today, grow with AI a little more.

Learn an AI skill every day.

Save a little time daily.

Improve a bit every day.

SasaDaily, growing with you.

Recommended Reading

AI Is Not Just on Screens: Japan Is Bringing Physical AI into Factories, Healthcare, and Everyday Life

Chinese Humanoid Robots Can Punch and Sprint, So Why Are They Harder to Use in Factories Than Traditional Robot Arms?

Factories Are Waiting for Humanoids, So Why Did Harmoni Just Stick a Tablet on an Old Machine?