{"slug": "xhigh-vs-high-one-effort-level-crushed-the-agent-testing-tool-by-61-points", "title": "xhigh vs high: One Effort Level Crushed the Agent Testing Tool by 61 Points", "summary": "A study of 90 independent coding-agent runs found that raising reasoning effort from 'high' to 'xhigh' increased first-try-perfect runs from 28 percent to 89 percent, a 61-point swing, while costing 9 to 29 percent more. The browser-based testing tool raised cost by 42 to 68 percent with no improvement in functional score. The study, titled 'Reasoning effort, not tool access, buys first-try reliability in…', suggests that adjusting the effort parameter is a low-cost, high-impact change for agent reliability.", "body_md": "Member-only story\n\n# xhigh vs high: One Effort Level Crushed the Agent Testing Tool by 61 Points\n\nNinety independent coding-agent runs, one specification, one rubric. The variable that moved the needle was not the browser-testing tool, not the design-oriented system prompt, and not the harness. It was a single string in the request body.\n\nRaising reasoning effort from `high`\n\nto `xhigh`\n\ntook first-try-perfect runs from 28 percent to 89 percent. That is a 61-point swing, and it cost between 9 and 29 percent more. In the same study, the browser-based testing tool that everyone bolts onto their agent raised cost by 42 to 68 percent and improved the functional score by nothing at all. Not on logic. Not even on the interface-visible criteria it was supposedly there to catch.\n\nI have been running agent harnesses in production for a year and I had never once swept the effort parameter. I tuned prompts. I added tools. I swapped models. The dial sitting directly on top of all of it, the one that costs a five-character edit to change, I left on default.\n\nHere is what the data says, what the vendor docs quietly admit, and the harness you can run this afternoon to find your own number.\n\n## The setup: 90 runs, one spec, 42 points\n\nThe study is *Reasoning effort, not tool access, buys first-try reliability in*…", "url": "https://wpnews.pro/news/xhigh-vs-high-one-effort-level-crushed-the-agent-testing-tool-by-61-points", "canonical_source": "https://pub.towardsai.net/xhigh-vs-high-one-effort-level-crushed-the-agent-testing-tool-by-61-points-2fa3bc284343?source=rss----98111c9905da---4", "published_at": "2026-08-05 04:09:00+00:00", "updated_at": "2026-08-05 04:22:12.810947+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-agents", "ai-research"], "entities": [], "alternates": {"html": "https://wpnews.pro/news/xhigh-vs-high-one-effort-level-crushed-the-agent-testing-tool-by-61-points", "markdown": "https://wpnews.pro/news/xhigh-vs-high-one-effort-level-crushed-the-agent-testing-tool-by-61-points.md", "text": "https://wpnews.pro/news/xhigh-vs-high-one-effort-level-crushed-the-agent-testing-tool-by-61-points.txt", "jsonld": "https://wpnews.pro/news/xhigh-vs-high-one-effort-level-crushed-the-agent-testing-tool-by-61-points.jsonld"}}