HomeAISkild AI and NVIDIA Launch S1 Model fo
AI

Skild AI and NVIDIA Launch S1 Model for Video Robot Training

Skild AI’s S1 foundation model learns multistep physical tasks from a single video demonstration without retraining.

WHAT YOU NEED TO KNOW
  • Skild AI launched S1, a foundation model that executes unseen physical tasks from a single video demonstration without retraining.
  • Skild reached a $100 million annual revenue run rate 10 months after its first commercial deployment across more than 60 partner programs.
  • Foxconn, NVIDIA, and Skild are deploying the system on dual-arm robots to assemble NVIDIA Blackwell server components.

Skild AI launched a robot foundation model called S1 that executes previously unseen, long-horizon physical tasks from a single video prompt, NVIDIA reported. The system uses in-context learning to translate demonstrated movements into robot actions without updating model weights or requiring task-specific post-training.

Researchers developed the model on NVIDIA computing infrastructure under a collaboration covering synthetic data generation, simulation, and real-world deployment. The launch comes as Skild reached a $100 million annual revenue run rate 10 months after its first commercial deployment, supported by more than 60 partnerships across manufacturing, logistics, inspection, security, and food preparation.

Video demonstrations

Operators record a video of a task and feed it directly into S1 as an input prompt. The model interprets the intent, identified objects, and action sequences, and then maps them to the physical robot. S1 handles tasks running up to 10 minutes, including potting plants, cooking pancakes, brewing pour-over coffee, and assembling component kits.

In a plant-potting trial, Skild moved from recording the video to autonomous physical execution in 11 minutes. Tests across new multistep tasks showed S1 achieved an average per-step success rate of roughly 66 percent, compared with 9 percent for an equivalent baseline AI system. Skild calculated that a single video demonstration replaces approximately 380 hands-on training runs, which typically require between 50 and 100 hours of manual data collection.

Factory assembly

Foxconn, NVIDIA, and Skild are deploying the software on dual-arm manipulators for high-precision assembly of NVIDIA Blackwell hardware. In one demonstrated sequence, a robot places a busbar and limit block, secures 16 screws, and adjusts to physical disturbances throughout the assembly cycle.

Training workflows rely on NVIDIA Cosmos world foundation models to generate structured descriptions from video, while Isaac Sim and Isaac Lab provide simulation powered by the Newton physics engine. Skild and NVIDIA are also developing GPU-accelerated contact solvers that will be released inside Newton, while NVIDIA TensorRT optimizes inference speeds for production robots.

Xentir Media
Xentir Media NewsroomSource-backed AI and technology coverage, drafted by Xentir's automated editorial system under fixed human-set rules. See our editorial policy and AI usage policy.
J
Jomon · Founder & EditorFounder and editor of Xentir Media. Sets the editorial rules the newsroom system runs under, and is accountable for its corrections. About Jomon · [email protected]
The Xentir Brief
The developments worth knowing — one useful email.
Get the Brief →