cd /news/robotics/sail-scaling-in-context-imitation-le… · home › topics › robotics › article
[ARTICLE · art-145428] src=sakana.ai ↗ pub= topic=robotics verified=true sentiment=↑ positive

SAIL: Scaling In-Context Imitation Learning

Sakana AI and the University of Tokyo introduced SAIL (Scaling In-Context Imitation Learning), a method that raises the average success rate of VLM-generated robot trajectories from 25% to 73% by scaling the search budget from one candidate to 45 across six simulated manipulation tasks, to be presented at IROS2026. SAIL uses a policy VLM conditioned on a few demonstrations, tests generated trajectories in a simulator, and has an evaluation VLM review the resulting video to identify where progress stalled, with Monte Carlo tree search exploring alternatives before only the selected trajectory is sent to the physical robot. The collaborators reported that robot trajectory generation benefits from test-time scaling, with additional computation letting the model test and refine proposed actions in simulation.

read2 min views11 publishedSep 27, 2026
SAIL: Scaling In-Context Imitation Learning
Image: Sakana (auto-discovered)

Introducing “Scaling In-Context Imitation Learning” (SAIL) to be presented at IROS2026. This work is a collaboration between Sakana AI and the University of Tokyo.

What does a robot need before it can tackle a new task?

Teaching a robot something new usually starts with collecting demonstrations and training a policy. But foundation models have already learned from vast amounts of images, text, and robotics-related data. We wanted to see how much of that knowledge we could draw out for robot control without changing the model itself.

Recent demonstrations suggest that GPT-6 Astra can operate physical robots alongside its general language and vision capabilities. Earlier work has also shown that LLMs/VLMs can generate entire sequences of robot movements from a few demonstrations.

However, a foundation model does not necessarily produce a reliable robot trajectory in a single generation. Performance depends on the context provided, and a small error in a movement target can cause the entire task to fail.

We propose SAIL, a method for more reliable VLM-based robot trajectory generation through test-time scaling.

SAIL uses a policy VLM as a robot trajectory generator, conditioned on a few successful demonstrations. It tests the generated trajectory in a simulator and uses an evaluation VLM to review the resulting video and identify where progress stalled. The policy VLM then uses this feedback to revise the trajectory, with Monte Carlo tree search (MCTS) exploring alternatives while refining promising candidates. Only the selected trajectory is sent to the physical robot.

Across six manipulation tasks in simulation, increasing the search budget from one candidate to 45 raised the average rate of finding a successful trajectory from 25% to 73%. We also evaluated SAIL on a physical robot. Our results suggest that robot trajectory generation can benefit from test-time scaling, with additional computation enabling the model to test and refine its proposed actions in simulation.

We think there is more to learn about what existing models can do with this kind of feedback, and how far those improvements carry over to physical robots.

── more in #robotics 4 stories · sorted by recency
── more on @sakana ai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/sail-scaling-in-cont…] indexed:0 read:2min 2026-09-27 · —