# Nvidia's AI Harness Pushed Claude Opus 5 From 30% to 100% on ARC-AGI-3

> Source: <https://startupfortune.com/nvidias-ai-harness-pushed-claude-opus-5-from-30-to-100-on-arc-agi-3/>
> Published: 2026-08-23 16:07:14+00:00

*Nvidia just showed that a smarter harness can beat a smarter model outright. Claude Opus 5 scored 30% alone on the ARC-AGI-3 benchmark, then hit a perfect 100% wrapped in a system Nvidia built itself.*

On August 21, Nvidia published research that should worry anyone betting the next leap in AI comes from bigger training runs alone. Researchers took Anthropic's Claude Opus 5, tested it cold on ARC-AGI-3, one of the hardest benchmarks in AI right now, and watched it score 30.2%. Then they wrapped the same model, completely unchanged, in a system Nvidia calls AVO. The score jumped to 100%.

No retraining. No new parameters. No extra compute thrown at the model itself. Nvidia just changed the scaffolding around it, and that alone turned a middling result into a perfect one.

ARC-AGI-3, built by the ARC Prize Foundation, drops an AI into a 2D game with no instructions, no tutorial, no example of what winning even looks like. The agent has to explore the environment, work out the objective, and plan a sequence of moves toward it, the way a person picks up an unfamiliar video game and just starts pressing buttons. Humans clear it close to 100% of the time. Frontier models, tested straight out of the box, have historically struggled to break double digits. That gap is the whole point of the benchmark: it's built to catch models that pattern-match rather than actually reason.

AVO wasn't built for this at all. Short for Agentic Variation Operators, it started out as a coding agent Nvidia designed to optimize CUDA GPU kernels. Its structure is fairly simple: a main agent inspects the current context, plans a move, carries it out and evaluates the result. A supervisor watches from above. When the main agent stalls or keeps repeating a failing approach, it steps in and redirects the agent toward a different strategy. Persistent memory carries prior attempts and reasoning forward, so the agent never starts from zero on the next turn. According to Nvidia's technical blog, that architecture, transplanted with no changes for a game-based task, cleared all 183 levels across 25 environments in the ARC-AGI-3 public set, using 6,624 actions in total.

[Vertiv stock loses $12 billion in a week as bond yields rattle AI bets](https://startupfortune.com/vertiv-stock-loses-12-billion-in-a-week-as-bond-yields-rattle-ai-bets/)

Vertiv Holdings shed about $12.3 billion in market value, roughly 12%, over the week ending August 21, 2026, as surging Treasury yields hit AI infrastructure stocks. The selloff follows a July revenue miss the company blamed on order timing, and now traders are waiting on Nvidia's August 26 earnings to decide whether the drop is profit-taking or... - [why AI infrastructure stocks fell in August 2026](https://startupfortune.com/vertiv-stock-loses-12-billion-in-a-week-as-bond-yields-rattle-ai-bets/) - [bond yields impact on data center equipment companies](https://startupfortune.com/vertiv-stock-loses-12-billion-in-a-week-as-bond-yields-rattle-ai-bets/)

Same model. Completely different outcome.

Nvidia isn't the only one who's found this. OpenAI ran a version of the same experiment on August 1, publishing a post titled "How enabling two settings tripled our scores on the ARC-AGI-3 benchmark." With the official harness, GPT-5.6 Sol scored 13.3% on the public set. Turn on two existing API settings, retained reasoning and compaction, and the score jumped to 38.3%, while output tokens dropped six-fold. OpenAI's own explanation was blunt: the official harness had been discarding the model's private reasoning after every action and truncating history at 175,000 characters, forcing the agent to rediscover a game's rules from scratch on every turn. Fix the plumbing, and the score triples. Still, nobody got anywhere near Nvidia's perfect run.

## The scaling story just got harder to sell

Here's the part that should sting inside Nvidia's own building. The company has told investors it expects global AI data center spending to hit $3 trillion to $4 trillion a year by 2030, a forecast resting largely on the idea that more compute keeps buying more capability. Its own researchers just showed that a smarter wrapper around a weaker configuration of a model can outperform a supposedly more capable one, without adding a single extra GPU-hour of training. That's an awkward argument to sit next to a company that sells the chips behind the brute-force approach.

It's awkward for benchmark credibility, too. If a research team can take the same underlying model from 30% to 100% just by changing memory management and adding a supervisor, a raw ARC-AGI-3 score doesn't really tell you what a model can do on its own. It tells you what a model can do wrapped in whatever harness someone bothered to build for it. Expect pressure on the ARC Prize Foundation and rival labs to standardize harness rules, or at minimum to report scores with and without heavy scaffolding, the way OpenAI already did on August 1.

So a bare score doesn't mean much any more.

For enterprise buyers trying to figure out where next year's AI budget goes, the message is plainer than the research paper makes it sound. Frankly, the model you pick may matter less than the system you build around it. Nvidia just spent its own research budget proving that point against its own sales pitch.

**Also read:** [Anthropic Cuts Claude Opus Prices in Half as Enterprises Balk at the Bill](https://startupfortune.com/anthropic-cuts-claude-opus-prices-in-half-as-enterprises-balk-at-the-bill/) • [ShipHero CEO Aaron Rubin Says Claude Mythos Beat His Pentest Firm for $10K](https://startupfortune.com/shiphero-ceo-aaron-rubin-says-claude-mythos-beat-his-pentest-firm-for-10k/) • [OpenAI's AI Models Broke Out and Hacked Hugging Face, So It Built Daybreak](https://startupfortune.com/openais-ai-models-broke-out-and-hacked-hugging-face-so-it-built-daybreak/)

[Nvidia Warns Hyperscalers Its AI Server Prices Are Jumping More Than 15%](https://startupfortune.com/nvidia-warns-hyperscalers-its-ai-server-prices-are-jumping-more-than-15/)

Nvidia has told hyperscalers and server makers that AI server prices are rising more than 15% as memory chip costs soar, according to Bloomberg. The increase reflects a DRAM and HBM shortage that has already pushed Nvidia's own next-generation rack costs toward $7.8 million apiece. - [Nvidia AI server prices increasing 2026](https://startupfortune.com/nvidia-warns-hyperscalers-its-ai-server-prices-are-jumping-more-than-15/) - [memory chip shortage driving hyperscaler costs](https://startupfortune.com/nvidia-warns-hyperscalers-its-ai-server-prices-are-jumping-more-than-15/)
