cd /news/artificial-intelligence/xhigh-vs-high-one-effort-level-crush… · home topics artificial-intelligence article
[ARTICLE · art-87158] src=pub.towardsai.net ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

xhigh vs high: One Effort Level Crushed the Agent Testing Tool by 61 Points

A study of 90 independent coding-agent runs found that raising reasoning effort from 'high' to 'xhigh' increased first-try-perfect runs from 28 percent to 89 percent, a 61-point swing, while costing 9 to 29 percent more. The browser-based testing tool raised cost by 42 to 68 percent with no improvement in functional score. The study, titled 'Reasoning effort, not tool access, buys first-try reliability in…', suggests that adjusting the effort parameter is a low-cost, high-impact change for agent reliability.

read1 min views1 publishedAug 5, 2026
xhigh vs high: One Effort Level Crushed the Agent Testing Tool by 61 Points
Image: Pub (auto-discovered)

Member-only story

Ninety independent coding-agent runs, one specification, one rubric. The variable that moved the needle was not the browser-testing tool, not the design-oriented system prompt, and not the harness. It was a single string in the request body.

Raising reasoning effort from high

to xhigh

took first-try-perfect runs from 28 percent to 89 percent. That is a 61-point swing, and it cost between 9 and 29 percent more. In the same study, the browser-based testing tool that everyone bolts onto their agent raised cost by 42 to 68 percent and improved the functional score by nothing at all. Not on logic. Not even on the interface-visible criteria it was supposedly there to catch.

I have been running agent harnesses in production for a year and I had never once swept the effort parameter. I tuned prompts. I added tools. I swapped models. The dial sitting directly on top of all of it, the one that costs a five-character edit to change, I left on default.

Here is what the data says, what the vendor docs quietly admit, and the harness you can run this afternoon to find your own number.

The setup: 90 runs, one spec, 42 points #

The study is Reasoning effort, not tool access, buys first-try reliability in

── more in #artificial-intelligence 4 stories · sorted by recency
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/xhigh-vs-high-one-ef…] indexed:0 read:1min 2026-08-05 ·