Nvidia Shows a Better AI Harness Beat a Smarter Model on ARC-AGI-3 Nvidia reported a 100.00 RHAE score on the ARC-AGI-3 public set using its Agentic Variation Operators (AVO) system, up from Claude Opus 5's 30.16% published score, by wrapping the same model family in a harness with persistent memory and supervision. Nvidia's AVO completed all 183 levels in 6,624 environment actions, about 12% fewer than VISTA's 7,542 actions with Claude Opus 5, showing that the agent's wrapper can matter as much as the model itself. Nvidia's AVO result makes one point hard to dodge: the agent around a model can matter as much as the model itself. Claude Opus 5 already looked strong on ARC-AGI-3. ARC Prize's results page put it at 30.16% on July 24, 2026, the best published model score on a benchmark that drops an AI into 2D game environments with no stated rules or goals. Then Nvidia put the same model family inside its Agentic Variation Operators system, or AVO, and reported a 100.00 RHAE score across the 25-environment public set on August 21. That's the sharp part. The model did not suddenly become a different model. Nvidia changed the system wrapped around it, and the public-set result moved from roughly 30% to a perfect score, with AVO completing all 183 levels in 6,624 environment actions. The wrapper did the work ARC-AGI-3 is not a normal question-and-answer test. The agent has to explore unfamiliar grids, work out what actions do, infer the objective and keep improving from feedback. If you forget what you already tried, you waste moves. If you keep chasing a dead end, you lose. That is exactly where a harness starts to matter. According to Nvidia's developer blog, AVO is built around persistent memory and supervision. The memory carries forward prior attempts and evaluation results, plus whatever state still matters. The supervisor watches for stagnation or repeated unproductive cycles and can push the main agent toward another path. In the ARC-AGI-3 setup, Nvidia said the model received exact 64 x 64 text-grid observations, not image tokens, and had to infer rules and goals through interaction. Cerebras Unveils CS-4 Chip It Claims Is 30 Times Faster Than Nvidia GPUs https://startupfortune.com/cerebras-unveils-cs-4-chip-it-claims-is-30-times-faster-than-nvidia-gpus/ Cerebras launched its CS-4 rack-scale AI system on August 18, claiming up to 30 times faster inference than Nvidia GPUs by fusing three wafer-scale chips into one rack. The launch lands just as rival Groq pivots away from its own chips into renting Nvidia GPU capacity, sharpening the divide over how to win the AI inference speed race. - Cerebras CS-4 chip performance compared to Nvidia GPUs https://startupfortune.com/cerebras-unveils-cs-4-chip-it-claims-is-30-times-faster-than-nvidia-gpus/ - AI inference speed improvements with Cerebras hardware https://startupfortune.com/cerebras-unveils-cs-4-chip-it-claims-is-30-times-faster-than-nvidia-gpus/ Be careful with the headline number. Nvidia itself says this was the ARC-AGI-3 public set, not the semi-private or fully private competition sets. It also says the comparison with other systems is not a controlled ablation, because the runs differ in backend, observation format, memory, context management and other implementation details. That caveat matters. Still, it doesn't drain the result of meaning. Nvidia compared AVO's 6,624 actions with 7,542 actions reported by VISTA using Claude Opus 5 on the same 183 public levels, roughly 12% fewer actions. AVO did not merely finish the set. It finished it with fewer moves in that cross-system comparison. The model race looks thinner now For the last two years, a lot of the AI market has talked as if the whole question is which model wins the leaderboard. Anthropic, OpenAI and Google push scores. Enterprise buyers look at charts. Investors ask which lab has the frontier model this quarter. Nvidia's result cuts straight through that habit. If a better harness can change the outcome this much, then buying the smartest model is only part of the job. You also have to ask what memory it has, what tools it can use, how it handles failure and whether anything is watching when it starts going in circles. Here's the thing: an agent is not just a model with a nicer interface. It is a model inside a working system. That creates a pricing problem for the model labs. If a company can pair a cheaper or less fashionable model with a better harness and still get useful long-horizon work, the premium attached to the top model becomes harder to defend on capability alone. TechCrunch also pointed to Databricks research showing that harness choice can swing agent costs by about 2x. That is not a rounding error for a company running thousands of agent tasks. Nvidia is not neutral here. It sells the chips that run the models, the harnesses and the long tests around them. By publishing AVO as a reference architecture, Nvidia is making a broader argument for an open agent stack where memory, runtime, tools and infrastructure can be tuned together. That argument serves Nvidia well, because every serious version of that stack still needs compute. The AI Boom Just Pushed Nvidia's RTX 5090 Past $4,900 on Nvidia's Own Store https://startupfortune.com/the-ai-boom-just-pushed-nvidias-rtx-5090-past-4900-on-nvidias-own-store/ The RTX 5090 now costs $4,929.99 on Nvidia's own store, more than double its $1,999 launch price, as Samsung, SK Hynix, and Micron divert memory chip capacity to AI data centers. MSI warns gaming hardware prices could rise another 15 to 30% before the year is out. - why is nvidia RTX 5090 so expensive now https://startupfortune.com/the-ai-boom-just-pushed-nvidias-rtx-5090-past-4900-on-nvidias-own-store/ - nvidia RTX 5090 price increase AI demand 2026 https://startupfortune.com/the-ai-boom-just-pushed-nvidias-rtx-5090-past-4900-on-nvidias-own-store/ The useful lesson for you is more direct. Don't treat benchmark scores as if they describe a model in isolation. ARC-AGI-3 is now showing something messier and more practical: the surrounding machinery can decide whether a model looks average, strong or suddenly unbeatable on a public test. Also read: Dutch Regulator Fines Uber Nearly $1 Billion Over Automated Driver Suspensions https://startupfortune.com/dutch-regulator-fines-uber-nearly-1-billion-over-automated-driver-suspensions/ • Reddit May Cut Off Google's AI Access When Their $60 Million Deal Expires https://startupfortune.com/reddit-may-cut-off-googles-ai-access-when-their-60-million-deal-expires/ • North v3 Brings Azure, AI, and Data Spend Into One Financial View https://startupfortune.com/north-v3-brings-azure-ai-and-data-spend-into-one-financial-view/