Its SAIL method raised simulated success from 25% to 73% with a 45-node search, but physical testing covered one task and six trials.
By [Ryan Merket](https://runtimewire.com/author/ryan-merket)
· Published
Primary source: [Sakana AI on X](https://x.com/SakanaAILabs/status/2104348705956774255)
Why it matters #
SAIL tests whether extra inference-time computation can improve robot control without retraining a model. Its simulation results are substantial, but one physical task and six trials leave the key deployment question unresolved: whether the gains hold across varied real-world environments.
Sakana AI says its SAIL method improves robot-generated action plans by testing and refining them in simulation before a robot executes one. The Tokyo AI company described the research in a September 27th post on X; the work, co-authored with the University of Tokyo, has been accepted to IROS 2026.
https://x.com/SakanaAILabs/status/2104348705956774255 The paper predates the post. Its first version appeared on arXiv on March 9th, and the authors revised it on September 19th. SAIL is a research method, not a commercial robot or product: it asks whether a vision-language model can control a robot more reliably by spending extra computation on candidate actions instead of retraining the model.
The approach reflects a practical question for Sakana AI CEO and co-founder David Ha (@hardmaru), who previously led Google Brain's Tokyo research team: how much can existing models do when researchers combine them with new systems around them? Sakana AI's company profile also identifies co-founder and CTO Llion Jones as a co-author of the Transformer paper "Attention Is All You Need." SAIL's authors are Makoto Sato and Yusuke Iwasawa of the University of Tokyo, and Yujin Tang and So Kuroki of Sakana AI. Sato conducted the work during an internship at Sakana AI.
SAIL starts with a vision-language model that receives a few successful demonstrations and generates a complete robot trajectory. The system runs that trajectory in a simulator, where a second model assesses task progress from video. Monte Carlo tree search uses that feedback to explore alternatives and revise promising plans. Only the selected trajectory is sent to the physical robot. In the experiments, Gemini Robotics-ER 1.5 served as both the policy model and evaluator, with no changes to its weights.
In six simulated manipulation tasks, the system found a successful trajectory in an average of 25% of cases with one search node, rising to 73% at 45 nodes. The benchmark covered 20 initial configurations for each task. Sakana AI's project page says success meant passing the simulator's ground-truth check within the search budget; the model's own estimates of progress guided the search but did not determine the reported result. That distinction makes the headline improvement a measure of whether the search found a working trajectory in simulation, rather than evidence of a comparable success rate in everyday robot use.
The physical test was much narrower. Using a LeRobot SO-101 arm for a block-placement task, the researchers reconstructed the scene from color and depth observations, searched for a trajectory in simulation, then ran the selected plan on the robot. With a budget of 15 candidate trajectories and randomized object positions, the method succeeded in five of six trials. The authors attributed the failure to pose-estimation error and differences in contact dynamics between the simulator and the real environment.
SAIL's own limits are consequential: the robot executes trajectories open-loop, without visual feedback while moving; search adds computation; and real-world validation covered only one task and six trials per method. The paper also reports that successful search trajectories could train an imitation policy; that approach likewise succeeded in five of six trials and reduced execution time compared with MCTS search. Those results point to a possible way to transfer expensive search into a faster policy, but the small trial count does not establish reliability across tasks or settings.
The bet is on inference-time work: use simulation, feedback and repeated attempts to extract more useful behavior from a model that has already been trained. Sakana AI's SAIL project page documents the method and experiments, while the paper reports success rates of up to 95% on complex simulated tasks. Those best-case results sit alongside the 73% average at the largest search budget and the limited physical evaluation. The research therefore offers a measurable result in simulation and an early hardware demonstration, not yet a broad claim about dependable robot autonomy.