Member-only story
Ninety independent coding-agent runs, one specification, one rubric. The variable that moved the needle was not the browser-testing tool, not the design-oriented system prompt, and not the harness. It was a single string in the request body.
Raising reasoning effort from high
to xhigh
took first-try-perfect runs from 28 percent to 89 percent. That is a 61-point swing, and it cost between 9 and 29 percent more. In the same study, the browser-based testing tool that everyone bolts onto their agent raised cost by 42 to 68 percent and improved the functional score by nothing at all. Not on logic. Not even on the interface-visible criteria it was supposedly there to catch.
I have been running agent harnesses in production for a year and I had never once swept the effort parameter. I tuned prompts. I added tools. I swapped models. The dial sitting directly on top of all of it, the one that costs a five-character edit to change, I left on default.
Here is what the data says, what the vendor docs quietly admit, and the harness you can run this afternoon to find your own number.
The setup: 90 runs, one spec, 42 points #
The study is Reasoning effort, not tool access, buys first-try reliability in…