# Do testing instructions improve coding agent correctness? What 30 conditions on one Zstd task measured

> Source: <https://stackness.dev/blog/do-testing-instructions-improve-coding-agent-correctness-what-30-conditions-on-one-zstd-task-measure>
> Published: 2026-10-10 18:56:38+00:00

# Do testing instructions improve coding agent correctness? What 30 conditions on one Zstd task measured

Mostly not. In [Dan Luu's eval](https://danluu.com/agentic-testing/) published on 7 September 2026, [OpenAI Codex](https://stackness.dev/tools/codex-cli) running [GPT-5.6 Sol](https://stackness.dev/tools/gpt) wrote a [Zstandard](https://stackness.dev/tools/zstd) decoder in [Rust](https://stackness.dev/tools/rust) about 4,800 times under 30 testing conditions, 80 runs per condition at each of two effort levels. Telling it to use test-driven development cut fully correct runs from 63.8% to 58.8% at medium effort and from 81.2% to 73.8% at xhigh. The best instruction beat no instruction by 11 points at most, which by my arithmetic is inside the noise of 80 runs.

## Do testing instructions improve coding agent correctness?

Testing instructions did not clearly improve coding agent correctness in Dan Luu's 30-condition eval. The Default condition, with no added instruction, ranked 9th of 30 on the mean of both effort levels, with 72.5% of runs passing all 37 hidden tests. Seven of the eight conditions above it led by 2.5 points or less. Only Luu's own five-bullet skill pulled clear, at 78.8%.

| Condition | Fully correct, medium | Fully correct, xhigh | Mean | Cost per run, medium | Cost per run, xhigh | 
|---|---|---|---|---|---|
| Luu's five-bullet skill | 73.8% | 83.8% | 78.8% | $3.99 | $5.86 | 
| [proptest](https://stackness.dev/tools/proptest) | 68.8% | 81.2% | 75.0% | $3.49 | $5.44 | 
| "Use property-based testing" | 66.2% | 83.8% | 75.0% | $3.51 | $5.16 | 
| "Audit and fuzz risky areas" | 75.0% | 73.8% | 74.4% | $4.07 | $4.98 | 
| [Kani](https://stackness.dev/tools/kani) | 72.5% | 76.2% | 74.3% | $4.23 | $6.70 | 
| "Make no mistakes" | 67.5% | 81.2% | 74.3% | $3.76 | $5.41 | 
| Audit first | 60.0% | 86.2% | 73.1% | $3.85 | $6.66 | 
| **Default, no instruction** | **63.8%** | **81.2%** | **72.5%** | **$3.31** | **$5.92** | 
| [Insta](https://stackness.dev/tools/insta) | 66.2% | 76.2% | 71.2% | $3.66 | $5.51 | 
| TDD | 58.8% | 73.8% | 66.3% | $3.27 | $5.50 | 
| [QuickCheck](https://stackness.dev/tools/quickcheck) | 57.5% | 75.0% | 66.2% | $3.14 | $5.66 | 
| [Lean 4](https://stackness.dev/tools/lean) | 57.5% | 73.8% | 65.7% | $2.82 | $4.80 | 
| [Hegel](https://stackness.dev/tools/hegel) skill | 55.0% | 76.2% | 65.6% | $4.34 | $8.06 | 
| [Verus](https://stackness.dev/tools/verus) | 51.2% | 75.0% | 63.1% | $3.60 | $4.93 | 

14 of the 30 conditions, from the data labels of the [interactive chart](https://danluu.com/interactive/agentic-testing/average-cost-vs-perfect-37.html) on Luu's page. The mean column is my arithmetic.

## TDD scored 5 to 7 points below no instruction

A test-driven development instruction made Codex write twice as many tests and fewer correct decoders. Under TDD, 58.8% of medium runs and 73.8% of xhigh runs passed every hidden test, against 63.8% and 81.2% with no instruction, at about the same cost per run. TDD ranked 24th of 30, and it also underperformed in Luu's second eval, on the IMAP RFC.

The agents followed the instruction: 67 of 160 TDD runs had failing tests before any substantial implementation, against 0 of 160 under Default. Few of those tests were fine-grained, and on the four-stream Huffman jump table TDD runs were more likely to skip the hard case, for example by making all four streams identical. Luu concludes that "Getting their own tests to pass more iteratively tended to get agents to write more incorrect tests that would enforce incorrect behavior."

Birgitta Böckeler's [smaller TDD comparison](https://martinfowler.com/articles/exploring-gen-ai/tdd-in-the-agent-loop.html) on martinfowler.com, run in August 2026 with [Claude Sonnet](https://stackness.dev/tools/claude-sonnet) 4.6 on three Python tasks and another model as the judge, found "no clearly discernable difference" in quality. TDD used 2.96 to 8.5 times the tokens, and agents "sometimes skipped or faked the red step". She has stopped asking agents to write tests first. The cost results differ because Luu counts dollars per run and Böckeler raw tokens, cache reads included.

## Property-based testing and a five-bullet skill led, inside the noise

The best testing instruction in Luu's eval depends on how you rank the results. His own five-bullet skill had the best mean of both effort levels, 78.8%. Proptest and a plain "use property-based testing" tied at 75.0%, the best of the 26 prompt addendums. "Audit first" posted the best single score, 86.2% at xhigh, at 12% more cost per run than no instruction.

In the property-based condition every agent picked proptest, so it was a second proptest run. Luu calls the property tests "mostly not very good", though proptest's shrinking "did sometimes provide some value". His skill, written in about two minutes and not ready for use by his own account, did not work as designed: agents almost never ran its fresh-context re-derivation step. Its first bullet asks the agent to "state likely mistakes and plausible alternative interpretations" for bug-prone areas before implementing.

"Make no mistakes", a placebo, scored 67.5% and 81.2%, level with or above Default. Luu argues that "a no-op is better than getting agents to do ineffective things".

With 80 runs a cell, one percentage carries roughly ±10 points of 95% uncertainty, so a gap between two conditions needs about 14 points to stand out. By my arithmetic, treating runs as independent, no condition clears that against Default. Luu cautions against "drawing any kind of strong conclusions from the ordering".

## The Hegel skill cost 41% more at xhigh and added nothing

The official Hegel property-testing skill was the clearest case of an instruction that burns tokens. Compared with the plain "use Hegel" prompt, it raised cost per run by 26% at medium and 41% at xhigh, from $5.73 to $8.06. Correctness stayed within noise: 55.0% against 58.8%, and 76.2% against 77.5%. It was the most expensive of the 30 conditions at both effort levels.

The skill is 34,000 characters, loads a 45,000-character Rust reference, and agents re-read it on many actions. The re-reading alone added 16% and 18% to the bill. A Rust testing skill from a collection Luu counted at 250,000 GitHub stars tells agents to do red-green TDD. The earlier a run read it, the more tests it added and the worse it did; the 16 runs that never opened it or read it late all reached 100%.

Formal methods were cheap but aimed at the wrong code. Lean 4 was the cheapest condition at both effort levels, and its proofs missed the bug-prone parts of the decoder. Kani was the only formal tool applied to the Zstd code itself, and it caught a non-trivial bug in 1 of 160 runs. Under [TLA+](https://stackness.dev/tools/tla-plus), 159 of 160 runs built a model, and Luu found no case where it changed the Rust code. The move of [writing TLA+ specifications alongside an agent](https://stackness.dev/moves/writing-tla-specifications-alongside-agentic-implementation) assumes the model feeds back into the code, which it did not do here.

## Why do agent-written test suites pass while asserting nothing?

Agent-written test suites pass while asserting nothing because the same agent writes the code and the check. The ExecCritic authors put it as "their errors can agree and create false confidence". On SWE-bench Verified, a [February 2026 study](https://arxiv.org/abs/2602.07900) counted 5.16 assertions against 25.00 value-printing statements per resolved task in the tests [Claude Opus](https://stackness.dev/tools/claude-opus) 4.5 wrote.

Luu's runs show the same pattern in other forms:

- **Verus:** agents "would frequently write vacuous proofs that were effectively A => A".
- **Differential testing:** none of the agents built two full implementations, and they "encoded the same bug in both versions".

A [coverage study of 4,882 agent pull requests](https://arxiv.org/abs/2607.18057) from Codex, [GitHub Copilot](https://stackness.dev/tools/github-copilot), [Cursor](https://stackness.dev/tools/cursor), [Claude Code](https://stackness.dev/tools/claude-code) and [Devin](https://stackness.dev/tools/devin), published 20 July 2026, found that agent tests added coverage in only 22.5% of Python and 35.9% of Java pull requests that included tests. In the Java pull requests that gained nothing, agents deleted 2.6 times as many tests as they added. Some of it is gaming: in [ImpossibleBench](https://arxiv.org/abs/2510.20270), Claude models from 2025 cheated mainly by modifying the test cases.

One guard is to fail the build when a suite asserts nothing. [PHPUnit](https://stackness.dev/tools/phpunit) exits 0 on zero-assertion tests by default, and [one project found 48 such methods](https://github.com/FJCF76/PromptingPress/issues/1110) before setting `failOnRisky="true"`.

## Testing gains show up when the tests come from outside the agent

Studies that found testing helped a coding agent gave it tests it did not write. [TDDev](https://arxiv.org/abs/2605.17242), a multi-agent TDD setup built on human-written acceptance tests, gained 15.5 to 23.7 points with capable models. [ClassEval-TDD](https://arxiv.org/abs/2602.03557) hands the model its public tests and gained 12 to 26 points. Luu and Böckeler asked the agent to write its own tests and found no gain.

[ExecCritic](https://arxiv.org/abs/2609.09133), published on 8 September 2026, measured the difference on SWE-bench Verified with a fixed repair agent:

- No tests: 61.2% resolved.
- Tests from a weak test agent: 57.3%.
- Tests from GPT-5.6 Sol: 65.3%.
- A trained test agent whose tests were frozen before the separate repair agent started: 72.6%. Compute was not matched.

A verification loop adds what a test instruction cannot: it changes who writes the check, makes the check fail on the old code first, and stops the coder from editing it. ImpossibleBench found the same lever, with read-only tests cutting cheating. One result points the other way: in Luu's eval, the 42 of 160 audit runs that handed review to an independent agent scored worse, which he notes may not be causal.

## Verification moves on Stackness sit in the build, not the prompt

The verification moves Stackness members record are build gates, not prompt instructions. As of 10 October 2026, four moves by real members are about verification, and three of them are mine. None of the four has an end date, and none tells the agent which test technique to use.

- [One `make check` target, and it is non-negotiable](https://stackness.dev/moves/one-make-check-target-and-it-is-non-negotiable)
- [E2E scenarios in markdown, test scripts as build artifact](https://stackness.dev/moves/e2e-scenarios-in-markdown-test-scripts-as-build-artifact)
- [Give the E2E agent its own isolated browser](https://stackness.dev/moves/give-the-e2e-agent-its-own-isolated-browser)
- qarep's markdown and XSS probe

I recorded my three on 22 August 2026, seven weeks before this post, and still run all three. The catalog move [instructing agents to use specific test techniques](https://stackness.dev/moves/instructing-agents-to-use-specific-test-techniques-before-shipping) was filed from the trend feeds on 22 September with Luu's post as a source.

Among the public profiles behind [the AI coding tools developers list](https://stackness.dev/categories/ai-tools), 7 list Claude Code and 2 list Codex. These are small numbers.

Outside Stackness, the evidence on what lasts is self-reported. After six months of AI-written [Playwright](https://stackness.dev/tools/playwright) tests, [Nilesh Raut](https://dev.to/speaklouder/i-let-ai-write-my-tests-for-6-months-here-is-what-actually-survived-production-4h2) still writes the test plan himself and rewrites "every assertion". Böckeler now checks regression quality with mutation testing.

## Method and sample

Luu's eval covers one task, one language and one harness: Codex with GPT-5.6 Sol writing a Zstd decoder in Rust from the RFC and its errata, single-turn, in a container without internet access. Each of the 30 conditions, 26 prompt addendums and 4 skills, ran 80 times at medium and 80 at xhigh effort. A run counted as correct only if it passed all 37 hidden tests.

What the numbers do not show:

- **Hidden tests:** agents wrote them, and Luu did not check them by hand. His earlier post on the same task mentions 34 tests, not 37.
- **Missing runs:** one formal-methods condition lost runs to out-of-memory errors, leaving 65 and 67.
- **Unpublished details:** the cost basis and the exact prompt text, except for Luu's skill. Max effort was run but not shown.
- **One task:** an IMAP eval gave results that "weren't materially different". Luu expects real-world failure modes to be "the same or worse", since RFC specs are clearer than most real work.

The Stackness counts come from public profiles as of 10 October 2026, with accounts marked as bots removed ([data sources](https://stackness.dev/about/data-sources)).

## Key numbers

- **58.8% vs 63.8%** of medium-effort runs fully correct with a TDD instruction and with none, and**73.8% vs 81.2%** at xhigh. Codex with GPT-5.6 Sol, 80 runs each,**7 September 2026** ,[Dan Luu](https://danluu.com/agentic-testing/) .
- **72.5%** , the no-instruction mean across both effort levels,**9th of 30** conditions. The best was**78.8%** , Luu's five-bullet skill,**7 September 2026** .
- **+26% and +41%** cost per run for the Hegel skill over the plain Hegel prompt, at medium and xhigh, with no correctness gain,**7 September 2026** .
- **61.2% to 57.3%** SWE-bench Verified resolve rate when a weak test agent wrote the tests; strong tests raised it to**65.3%** ,**8 September 2026** ,[ExecCritic](https://arxiv.org/abs/2609.09133) .
- **5.16 vs 25.00** assertions against value prints per resolved task in Claude Opus 4.5's tests,**8 February 2026** ,[Chen et al.](https://arxiv.org/abs/2602.07900)
- **22.5% and 35.9%** of Python and Java agent pull requests with tests whose tests added coverage,**20 July 2026** ,[Dipongkor et al.](https://arxiv.org/abs/2607.18057)
- **4** verification moves by real Stackness members,**none** retired, as of**10 October 2026** ,[Stackness](https://stackness.dev/about/data-sources) .

## Tools in this post

## Use any of these tools?

Put them on a Stackness profile, say how you use each one and see who pairs them the same way. It takes a couple of minutes.

[Show my stack](https://stackness.dev/register)
