Skild AI launched a robot foundation model called S1 that executes previously unseen, long-horizon physical tasks from a single video prompt, NVIDIA reported. The system uses in-context learning to translate demonstrated movements into robot actions without updating model weights or requiring task-specific post-training.
Researchers developed the model on NVIDIA computing infrastructure under a collaboration covering synthetic data generation, simulation, and real-world deployment. The launch comes as Skild reached a $100 million annual revenue run rate 10 months after its first commercial deployment, supported by more than 60 partnerships across manufacturing, logistics, inspection, security, and food preparation.
Video demonstrations
Operators record a video of a task and feed it directly into S1 as an input prompt. The model interprets the intent, identified objects, and action sequences, and then maps them to the physical robot. S1 handles tasks running up to 10 minutes, including potting plants, cooking pancakes, brewing pour-over coffee, and assembling component kits.
In a plant-potting trial, Skild moved from recording the video to autonomous physical execution in 11 minutes. Tests across new multistep tasks showed S1 achieved an average per-step success rate of roughly 66 percent, compared with 9 percent for an equivalent baseline AI system. Skild calculated that a single video demonstration replaces approximately 380 hands-on training runs, which typically require between 50 and 100 hours of manual data collection.
Factory assembly
Foxconn, NVIDIA, and Skild are deploying the software on dual-arm manipulators for high-precision assembly of NVIDIA Blackwell hardware. In one demonstrated sequence, a robot places a busbar and limit block, secures 16 screws, and adjusts to physical disturbances throughout the assembly cycle.
Training workflows rely on NVIDIA Cosmos world foundation models to generate structured descriptions from video, while Isaac Sim and Isaac Lab provide simulation powered by the Newton physics engine. Skild and NVIDIA are also developing GPU-accelerated contact solvers that will be released inside Newton, while NVIDIA TensorRT optimizes inference speeds for production robots.
