Skild AI unveiled its S1 robot foundation model that can learn a task from a single human‑task video, enabling multi‑step jobs without retraining.
Skild AI’s new S1 foundation model can teach a robot to complete a ten‑minute, multi‑step task after watching just one video of a human performing the same job.
How the S1 Model Works
The S1 model processes a single video clip of a person executing a task, extracts the sequence of actions, and translates them into robot‑compatible commands. Unlike traditional robot training that requires thousands of demonstrations, S1 leverages a large‑scale transformer architecture to generalize from that lone example.
Key Capabilities
- Learn multi‑step procedures without additional data collection
- Adapt to variations in object placement and orientation
- Operate on standard robotic hardware without custom firmware
In tests, the model successfully guided a robot arm to assemble a simple kitchen gadget, sort items on a conveyor, and perform a basic cleaning routine—all after a single human demonstration.
Implications for Industry
If the approach scales, manufacturers could reduce the time and cost of programming robots for new products, while small businesses might deploy automation without hiring specialist engineers.
Skild AI plans to release an API that lets developers upload a video and receive a ready‑to‑run robot program, aiming to make the technology accessible across logistics, retail, and home‑assistant markets.
For more details, see Superpower Daily coverage of Skild AI’s S1 model.