Robotics startup Skild AI has launched its S1 robot foundation model, developed on NVIDIA AI infrastructure to enable robots to learn unseen, multi-step tasks from a single video demonstration. Using in-context learning, the model interprets video prompts and translates them into physical actions without updating its weights or requiring task-specific post-training. Skild AI also reported reaching a $100 million annual revenue run rate 10 months after its initial commercial deployment.

The S1 model can perform unfamiliar physical tasks lasting up to 10 minutes—such as kit assembly, pour-over coffee brewing, and plant potting—achieving a 66% step success rate in tests compared to 9% for baseline systems. Skild estimates one video example provides learning value equivalent to roughly 380 manual training runs. Skild, NVIDIA, and Foxconn are currently deploying the technology on dual-arm manipulators for high-precision assembly of NVIDIA Blackwell systems.

Why it matters

  • In-context learning via video prompts drastically reduces robotics deployment cycles by removing the need for task-specific dataset collection and fine-tuning.

  • Adaptable physical AI models enable dynamic factory retooling, allowing hardware operators to introduce new processes without costly reprogramming.

  • Rapid commercial traction in physical AI signals maturing enterprise adoption for flexible, non-preprogrammed robotic foundation models.

Source: blogs.nvidia.com