Nvidia’s AVO agent completes ARC-AGI-3 benchmark with 100% success rate Nvidia's AVO agent, powered by Anthropic's Claude Opus 5, achieved a 100% success rate on the ARC-AGI-3 benchmark, completing all 183 levels across 25 environments on August 21 while using 12% fewer actions than the previous best, VISTA. The system-level architecture, which includes persistent memory and iterative loops, boosted Claude Opus 5's standalone performance from roughly 30.2% to a perfect score, though results were only disclosed on public datasets. Via nvidia.com Nvidia’s AVO agent completes ARC-AGI-3 benchmark with 100% success rate The agentic system completed all 183 levels across 25 environments while using 12% fewer actions than the previous best, turning Claude Opus 5's baseline performance from roughly 30% to a perfect score. Nvidia just did something no AI system has managed before: a perfect score on ARC-AGI-3, the interactive reasoning benchmark designed to test whether AI agents can figure out unfamiliar environments without anyone holding their hand. The company’s AVO system, short for Agentic Variation Operators, completed all 183 levels across 25 public environments on August 21. It did so while requiring 12% fewer actions than VISTA, the previous top performer. The underlying model powering AVO is Anthropic’s Claude Opus 5, which on its own manages roughly 30% on the same benchmark. Nvidia’s system-level architecture turned that into 100%. What ARC-AGI-3 actually tests ARC-AGI-3 isn’t your typical AI benchmark where a model answers multiple-choice questions or generates text. It’s closer to dropping an agent into a video game it’s never played and asking it to figure out the rules, objectives, and win conditions entirely on its own. The benchmark requires agents to discover patterns, navigate unfamiliar environments, and make decisions without explicit instructions. The 25 environments and 183 levels are designed to expose whether an AI system can genuinely reason and adapt, or whether it’s just pattern-matching against training data. Claude Opus 5, running standalone, tops out at about 30.2% on these tasks. That gap between 30% and 100% is entirely attributable to the engineering wrapped around the model, not the model itself. How AVO bridges the gap The architecture Nvidia built around Claude Opus 5 is where the real story lives. AVO uses persistent memory, meaning it retains what it learns across interactions rather than starting fresh each time. It runs iterative loops that inspect the current state, plan the next move, implement it, and evaluate the results before cycling again. There’s also a supervisory system that monitors progress and helps the agent course-correct when it gets stuck. Before the ARC-AGI-3 results, AVO had already shown its chops in more practical settings. In autonomous GPU-kernel optimization tasks running on Nvidia’s DGX B200 hardware, the system outperformed cuDNN by up to 3.5% and exceeded FlashAttention-4’s performance by up to 10.5%. What this means for AI development Nvidia is well-positioned here regardless. The company already dominates AI hardware sales, and demonstrating that its own agentic systems can dramatically amplify model performance adds a software layer to its competitive moat. That said, some caveats apply. Nvidia disclosed results only on public datasets, not private ones. The ARC-AGI benchmark has historically used private test sets to prevent overfitting, where systems effectively memorize answers rather than genuinely reasoning. Without private set results, the 100% score is impressive but incomplete as a measure of general reasoning ability. There’s also no word on commercialization. AVO exists as a research demonstration for now, not a product you can buy or integrate. Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy https://cryptobriefing.com/editorial-policy/ .