By Ambi Robotics
Blog #
Using a new Graph-as-Policy-based Agentic Robotics harness within AmbiOS, AI agents solved in 10 hours a package-placement problem that would have taken engineers several weeks to address. The solution is now deployed across 30% of Ambi Robotics’ U.S. fleet, including systems operating for a Fortune 50 package shipping company. To the best of our knowledge, this is the first time Agentic Robotics has solved a real production problem.
Improving performance on commercially-deployed robots requires significant human effort. Engineers must review key data offline (e.g. videos, images, sensor data, and logs), hypothesize root causes, and test updates. This process is slow and requires significant domain knowledge.
“Agentic Robotics” offers an exciting new approach. The idea is to use AI coding agents to automatically analyze data, identify problems, and solve them to improve robot performance. This combines rapid advances in AI coding with classical model-based robotics modules and learned model-free policies. This requires agentic robotics “harnesses” to guide AI coding systems to solve the real robotics problems.
Several new research papers, including Code-as-Policy (CaP) [1, 2] and Graph-as-Policy (GaP) [3], show surprising promise for rapidly generating reliable robot control systems with agentic harnesses that constrain agents to use libraries of relevant robot skills.
Building on these ideas, we developed the AmbiOS Agentic Robotics Harness. It builds on GaP to update computation graphs that are composed of existing robot skills. The harness gives agents access to two key pieces of infrastructure. First, agents access datasets of real robot events from the cloud that can be filtered by a set of known success and failure modes. Second, agents use a simulation environment to replay these events with new self-generated behaviors and test graph edits, evaluating the outcomes in the simulation. The harness constrains how agents work, ensuring that the final result can be executed without compromising safety checks and error handling.
Figure 1. AmbiOS Agentic Robotics harness architecture for performance improvements on real production robots.
Task #
The AmbiSort robot picks, scans, and places random packages with a robot arm and a gantry receives the package and sorts it into a bag based on data like the zip code. The target bags have a fixed opening size through which items must fit to fall into the bag without jamming.
We applied the AmbiOS Agentic Robotics Harness to a problem where subtle edge cases impact critical operational metrics: throughput, sort accuracy, and uptime. Millimeter-level inaccuracies can have a significant effect on throughput.
Packages that are on the boundary of fitting into the bag opening are challenging because the robot arm cannot estimate item dimensions and pose with perfect precision due to sensor noise and item deformation. Furthermore, the robot prefers placing items with the short side going into the bag to maximize space utilization. As a result, these edge case items are too wide to fit through the bag opening, forcing the system to reject and retry packages, losing valuable time.
We estimated that this issue was reducing capacity by approximately 31k sorts per year per robot for a typical facility running 12 hours a day, 6 days a week. Fixing this issue with hand-tuning would take several weeks of engineering time.
Figure 2. Example of a retry due to the item not fitting into the sack. The item is drooping slightly when lifted, leading the robot to underestimate the item width and place it in a poor orientation. The gantry drops the item back into the input bin to retry, costing valuable time.
We used the AmbiOS Agentic Robotics Harness with Anthropic’s Claude models including both Sonnet 4.6 and Opus 4.8. The agents analyzed images and data from 500 events over the previous week of production, 50% of which were the known failure mode and 50% which were selected at random to avoid regressions. Within 10 hours, the agents generated and evaluated three solution hypotheses.
Results #
We deployed these three hypotheses into real production A/B tests in customer warehouses. Each solution was compared to a baseline software version. Here are the results after 10 days:
| Hypothesis | Estimated Change in Sorts per Year per Robot |
|---|---|
| H1 | −374 |
| H2 | −1,497 |
| H3 | +15,725 |
Table 1. Comparison of throughput and estimated packages per year across three A/B test solutions deployed to real production robots.
For H3, the agent noticed that for large, dangling bags and paper mailers, the placement plan needs to carefully trade off orienting the item to fit in the destination bag and avoiding placements where the item could get stuck in grates on the placement platform. It optimized this tradeoff and decided to route more packages through a subgraph for planning dangling item placements because it results in a more predictable final orientation. In addition, the agent changed the flat item subgraph to only optimize for the final orientation of the item because it is easier to predict the final resting position of flat items than for bags. In real production A/B tests, H3 increased throughput by 4.2 packages per hour. Over 12 hours per day and 6 days a week (standard operation), this corresponds to over 15k packages per year per robot.
Figure 3. Placement behavior before and after S3, for a set of near-identical packages in production. Select an example to load it, then drag a view to rotate both.
Conclusion #
To the best of our knowledge, this is the first time that Agentic Robotics has solved a real production problem. The solution runs on 30% of the Ambi Robotics fleet sorting real packages with 24×7 availability. We are now applying the AmbiOS Agentic Robotics Harness to a number of other production problems, accelerating performance improvements.
Want to scale physical AI in real world deployments? We’re hiring!
CLICK HERE to see open roles.
References #
[1] Liang, Jacky, et al. “Code as policies: Language model programs for embodied control.” 2023 IEEE International conference on robotics and automation (ICRA). IEEE, 2023.
[2] Fu, Letian, et al. “CaP-X: A framework for benchmarking and improving coding agents for robot manipulation.” arXiv preprint arXiv:2603.22435 (2026). International Conference on Machine Learning (ICML), July 2026.
[3] Chen, Kaiyuan, et al. “GaP: A Graph-as-Policy Multi-Agent Self-Learning Harness For Variational Automation Tasks.” arXiv preprint arXiv:2607.05369. To Appear: Conference on Robot Learning (CoRL), Nov 2026.
Blog | 08.26.2026CARGO: Physical AI for Industrial Package StackingIntroducing CARGO (Contact-Aware Reinforcement-learned Generalized Object-stacking): Sim2Real reinforcement learning for package stacking #
Blog | 07.22.2026Physical AI: Why Form Factor MattersThe shape of a robot matters more than you think. For operations leaders evaluating robotic palletizing systems, learn how workspace geometry, and the shape of a robot's reach, affects pallet density, uptime, and throughput per square foot. #
Blog | 07.15.2026Gantry vs. 6-Axis Robots for Palletizing: How Workspace Geometry Affects Throughput, Uptime, and Floor SpaceThe shape of a robot matters more than you think. For operations leaders evaluating robotic palletizing systems, learn how workspace geometry, and the shape of a robot's reach, affects pallet density, uptime, and throughput per square foot.