{"slug": "nvidia-researchers-improve-ai-agent-reliability-with-a-judging-model", "title": "Nvidia researchers improve AI agent reliability with a judging model", "summary": "Nvidia researchers reported that their Mid-Harness method, which uses a separate verifier model to score multiple candidate actions before a terminal agent executes one, raised first-try success on the TerminalBench-Lite benchmark from 50.00% to 68.03% Pass@1 with eight sampled actions per step. The paper, published December 16, 2026 by authors including Minki Kang and Ehsan Hosseini-Asl, found the strongest results came when the TMAX-9B generator also served as its own verifier, and that action scaling beat trajectory scaling at lower token cost. The 68.03% Pass@1 rate still leaves roughly one task in three failing on the first attempt, and the results come from controlled benchmark settings.", "body_md": "# Nvidia researchers improve AI agent reliability with a judging model\n\nA new method called Mid-Harness has a verifier model pick the best command before a terminal agent runs anything\n\n[Nvidia](https://cryptobriefing.com/markets/nvidia/) researchers have found a fairly human fix for unreliable AI agents: think before you type. Their new method, called Mid-Harness, has an agent generate several possible actions, then lets a separate judge model choose which one actually runs.\n\nOn one benchmark, that extra moment of deliberation lifted the first-try success rate from 50.00% to 68.03%. For software that operates inside a command line, where one bad command can sink an entire task, that matters.\n\n## How Mid-Harness works\n\nMid-Harness has a generator model sample multiple candidate actions at each step, and a second model, the verifier, scores them. Only the top choice is forwarded for execution.\n\nThe target is what researchers call terminal agents. These are AI systems that work in command-line interfaces or call external tools, typing real commands into real environments.\n\nThose environments are stochastic, meaning outcomes can vary unpredictably from one run to the next. An agent might usually know the right command and still fumble it at the worst possible moment.\n\nThe paper frames this as a gap between generation and reliable execution. A model can be capable of producing a useful command without consistently producing it when it counts.\n\nMid-Harness aims to close that gap without retraining the generator. It simply adds a filter between thinking and doing.\n\n### AI, tech, and the markets they move—in one daily briefing.\n\nDaily. Free. Join 34,000+ readers across crypto, finance, and policy.\n\n## The numbers on TerminalBench-Lite\n\nThe team tested the method on TerminalBench-Lite, a benchmark for terminal agent tasks. Their headline result paired a TMAX-9B generator with GPT-5.6 Sol acting as the verifier.\n\nWith eight sampled actions per step, Pass@1 climbed from 50.00% to 68.03%. Pass@1 measures how often the agent succeeds on its first and only attempt, which is the scenario that matters most in production.\n\nAccording to the research findings, the most effective results emerged when TMAX-9B served as both generator and verifier, grading its own homework. It suggests a model may be better at recognizing a good command than reliably producing one on the first try.\n\nThe researchers also compared two strategies for spending extra compute. Action scaling, the Mid-Harness approach, samples many options at each individual step. Trajectory scaling instead runs entire task attempts multiple times and picks the best overall run.\n\nAccording to the findings, action scaling delivered better outcomes at a lower token cost than trajectory scaling alone. The two approaches also worked well together. The research reports consistent improvements across diverse models and benchmarks, not just the single headline pairing.\n\n## Background: Nvidia’s push on agent reliability\n\nThe paper was published on December 16, 2026, by authors including Minki Kang and Ehsan Hosseini-Asl of Nvidia. Earlier in 2026, Nvidia introduced ACES, a framework published around August 2026 for measuring the real runtime impact of agent skills. In September 2026, it followed with the Open Agent Safety Platform.\n\n## What this means\n\nIf action-level verification beats trajectory-level retries on token efficiency, as the findings indicate, it offers a cheaper path to better reliability. Checking each step costs less than redoing the whole job.\n\nIf a 9-billion-parameter model can effectively judge its own candidate actions, teams may not need to pay for a second, larger model to act as the referee, simplifying deployment and keeping costs down.\n\nThere are caveats. A 68.03% Pass@1 rate is a real improvement, but it still means roughly one task in three fails on the first attempt. Benchmarks are also controlled settings, and how well Mid-Harness holds up across messier real-world environments, longer tasks and different toolchains remains to be tested.\n\n**Disclosure:** This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our\n\n[Editorial Policy](https://cryptobriefing.com/editorial-policy/).", "url": "https://wpnews.pro/news/nvidia-researchers-improve-ai-agent-reliability-with-a-judging-model", "canonical_source": "https://cryptobriefing.com/nvidia-mid-harness-ai-agent-reliability/", "published_at": "2026-10-01 17:56:53+00:00", "updated_at": "2026-10-01 18:18:59.020553+00:00", "lang": "en", "topics": ["ai-agents", "ai-research", "artificial-intelligence", "large-language-models", "ai-safety"], "entities": ["Nvidia", "Mid-Harness", "TerminalBench-Lite", "TMAX-9B", "GPT-5.6 Sol", "Minki Kang", "Ehsan Hosseini-Asl", "ACES"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/nvidia-researchers-improve-ai-agent-reliability-with-a-judging-model", "markdown": "https://wpnews.pro/news/nvidia-researchers-improve-ai-agent-reliability-with-a-judging-model.md", "text": "https://wpnews.pro/news/nvidia-researchers-improve-ai-agent-reliability-with-a-judging-model.txt", "jsonld": "https://wpnews.pro/news/nvidia-researchers-improve-ai-agent-reliability-with-a-judging-model.jsonld"}}