cd /news/ai-agents/tdd-with-coding-agents-write-the-rul… Β· home β€Ί topics β€Ί ai-agents β€Ί article
[ARTICLE Β· art-142445] src=dev.to β†— pub= topic=ai-agents verified=true sentiment=Β· neutral

TDD With Coding Agents: Write the Rules, Then Check They Held

A developer working with IBM Bob as a coding agent found that tests written by agents often pass regardless of whether the code works, and proposes a set of instruction-file rules plus a post-green check to catch non-discriminating tests. The work, reproducible via a sample GitHub project with six test shapes and a line-by-line mutation runner, argues that existing TDD tooling such as TDD Guard enforces test sequence but not whether a test can actually fail. The author notes the patterns appeared across languages, codebases and tools, including instruction mechanisms like AGENTS.md, .bob/rules, CLAUDE.md and .cursor/rules.

by read14 min views2 publishedSep 30, 2026

Test-driven development and coding agents fit together unusually well. A test

is a precise specification the agent can check its own work against. The

red-green-refactor cycle keeps each change small enough to review. And the

"run the tests, read the output, try again" rhythm is exactly what these tools

are good at.

The guidance on how to do it has matured quickly, and most of it is sound. But

after working this way across different codebases and languages, with the agent

writing both the tests and the code, I kept finding tests that passed whether

the code worked or not.

Most of those were preventable. A handful of rules, written once into the

agent's instructions, stop the common shapes before they are written. The rest

needed a check after green, because no rule anticipates every way a test can

fail to discriminate.

This post is both: the rules worth writing down, and the thirty-second check

for what they miss.

The work behind this post used IBM Bob as the coding agent. Nothing here is

specific to it. Bob routes tasks across several frontier models rather than a

single fixed one, so a single session may not even have used the same model

throughout β€” and the same patterns turned up regardless of language, codebase

or tool.

Everything in this post is reproducible: sample project on GitHub β€” six test shapes and a runner that mutates one line at a time.

It's worth tracing how the advice has developed, because each stage solved a

real problem.

Keep the tests in human hands. The early guidance was that if a model

writes both the code and the tests, the same assumption ends up in both and

they agree with each other. So humans wrote the tests and the AI wrote the

code. That division worked, and it's still good advice when the spec matters

more than the speed.

Write the discipline down. Agents can produce twenty tests in the time you

write one, and most teams took that trade. The guidance shifted to enforcing

the cycle through whatever instruction mechanism the tool offers β€” an

AGENTS.md in the project root now works across several of them, alongside

tool-specific rules files like .bob/rules for IBM Bob, CLAUDE.md for Claude

Code and .cursor/rules for Cursor β€” carrying the project's conventions and

the test-first rule, one behaviour per cycle, a human gate before

implementing, and, in the stronger versions, committing the red test so the

failing state lives in git history.

Kent Beck, who originated TDD, has published a system prompt along these lines: always follow the cycle, write the simplest failing test first, implement the minimum needed to pass.

Enforce it mechanically. The newest approach doesn't rely on the agent

following instructions at all. TDD Guard hooks every file write and blocks it

unless the process was honoured: a failing test exists, you're on one test at

a time, you're working outside-in. Under the covers it spins up a second model

as a judge on each edit, because "did this follow TDD?" is a fuzzy question

that is easier to ask a model than to encode in rules.

Each stage is a genuine improvement, and the progression is the right one.

Prompting alone tends to produce what one practitioner aptly called "test

first, not test-driven" β€” all the tests, then all the code, in two large

steps. Instruction files make the cycle stick more often. Hooks make it stick unless

the judge misses something.

Every tool I looked at governs sequence: was the test written first, is it

one behaviour, did red precede green. That's the hard part to enforce, and they

enforce it well.

The step I'm suggesting governs something different: whether the resulting test can fail at all.

These go in whatever instruction mechanism your tool offers, and they are the

higher-value half of this post. Written once, they make most of the shapes in

the next section much less likely, though no instruction is followed perfectly:

Those rules prevent most of the shapes below. What follows is for the cases

they cannot anticipate.

After green, before moving on:

Break one production line on purpose. Say in advance which test should fail.

Run it. Confirm it fails on an assertion. Then restore the line.

That's mutation testing, done by hand, one line at a time, at review time

rather than in CI. Note the direction: the code is correct, you introduce a

deliberate defect, and the test failing is the good outcome β€” it's the test

doing its job.

A test that can't fail isn't a new problem, and it isn't specific to AI.

Mutation testing has existed for decades precisely because of it, and most TDD

guides mention it β€” usually a line in a metrics section beside coverage

thresholds, pointing at Stryker or mutmut. That framing makes it sound like a

quarterly exercise. Used per change, it's smaller and more immediate: the one

question that separates a test from a decoration.

What's different with agents is that TDD normally guards against this, and the

guard gets weaker. The red step is meant to be the proof β€” you watched the test

fail, so it can fail. That holds when a person writes one test and sees it go

red for the reason they expected. It holds less well when an agent produces the

test: the red often comes from the function not existing yet rather than from

the assertion, and the volume means few of them get inspected closely.

Better instructions narrow this a long way, and you should write them. But

instructions produce better tests; they do not produce proof that a given test

can fail. That distinction is the whole reason the check exists.

So the mutation asks what the red phase no longer reliably answers. A red test

proves it failed before the code existed β€” when everything failed. This asks:

now that the code exists, would this test notice if it broke?

Most of the time the answer is yes, it takes thirty seconds, and you move on.

Occasionally it isn't, and those are the cases worth writing about.

Two agent-written tests for the same behaviour: the service forwards the

caller's correlation ID to a downstream supplier. Both green.

Break the one production line they exist to protect:

- new SupplierClient(supplierTransport, { forwardCorrelation: true });
+ new SupplierClient(supplierTransport, { forwardCorrelation: false });

The service no longer forwards the header. It still compiles, which matters β€”

a mutation that doesn't compile tells you nothing. Then run the tests:

βœ“ hollow: correlation ID reaches the supplier
Γ— fixed:  correlation ID reaches the supplier

AssertionError: expected undefined to be 'abc'
  41|   expect(supplierSide.callCount).toBe(1);
  42|   expect(supplierSide.last()?.headers[HEADER]).toBe("abc");

Restore the line, confirm green, move on. Total cost: about thirty seconds.

The reasoning behind it: two passing tests tell you the tests and the code

agree. That's true when the code is right β€” and equally true when the test

can't tell the difference. From a green suite, those two situations look

identical.

Introducing a defect separates them. A test that genuinely checks the behaviour

has to notice, because its result depends on that behaviour. The hollow one

passed with forwarding switched on and with it switched off, which means its

result never depended on forwarding at all.

I've rebuilt each of these on a small fictional service in TypeScript so you

can run them (link at the end). Each has a hollow version that passes and

a fixed version that also passes. The difference only shows under mutation.

None of these come from carelessness. They're the kind of test a competent

developer writes and a reviewer approves.

Four of the six are preventable by a rule. Two are not, and those are the ones

that justify the check. I've marked each.

Preventable by rule 5.

The service should forward a caller's correlation ID to a downstream supplier.

The test records outbound requests and checks for the header:

const shared = new Recorder();
const service = buildService(new Transport(shared));
const caller = new Transport(shared, (req) => service.handle(req)); // same recorder

caller.send({ path: "/orders", headers: { [HEADER]: "abc" }, body: ORDER });

expect(shared.requests.some((r) => r.headers[HEADER] === "abc")).toBe(true);

The test's own client and the code under test share one recorder. The inbound

request already carries the header, so the recorder always sees it β€” whether

or not the service forwarded anything. Break the forwarding and the test stays

green.

The assertion is fine. The fixture defeats it. The fix is separate recorders

and an assertion on the side that matters:

expect(supplierSide.callCount).toBe(1);
expect(supplierSide.last()?.headers[HEADER]).toBe("abc");

This is the one I'd least expect to find by reading. The fixture looks

correct; only the mutation shows it isn't.

Preventable by rule 4.

The test constructs the objects by hand, with the right settings, instead of

using the factory production uses. When the factory stops passing the setting,

the test doesn't notice β€” it never calls the factory.

Convincing in review: real code, real requests, real assertions. Just not the

code that ships.

Preventable by rule 2.

expect(rec.requests.some((r) => HEADER in r.headers)).toBe(false);

some() over an empty array is false, so if the code path never runs, this

passes. "Nothing wrong happened" and "nothing happened" look identical. One

line fixes it: prove the calls were made first.

Preventable by rule 3.

An audit trail always contains the inbound entry, so asserting it isn't empty

can't tell you whether the reservation was recorded. Assert the specific entry.

Not preventable by a rule. Nothing in the test looks wrong; the argument for removing the assertion is reasonable until a mutation answers it.

A log's sequence number must advance only after a write succeeds. After a

failed write, the test checks the store is empty β€” which it always is, because

the write is all-or-nothing.

There's a reasonable argument for dropping the second assertion, that the

counter hasn't moved: the store is empty, so what could be wrong? The mutation

answers it. Advance the counter before the write, and the store is still

empty, but:

expected 2 to deeply equal +0

The counter has moved past entries that were never written. Only the

"redundant" assertion sees it.

Not preventable by a rule. The test is correct; the behaviour is invisible from where it is looking, and you only find that out by breaking the code.

Only the first request of an order should carry a parent ID. The integration

test checks what the downstream system recorded β€” but that system keeps the

value from the first call and ignores it afterwards. Sending it every time

changes nothing observable, and every integration test passes either way.

The integration test isn't wrong. It's looking where the behaviour is

invisible. A unit test on the outbound requests can see it.

The first instinct is to strengthen the assertion. That's usually the wrong

move, and it cost me two rewrites before I stopped reaching for it. A surviving

mutation means the test's result didn't depend on the behaviour β€” so the

question is why not, and the answer is often somewhere other than the

assertion.

Four things to check, in order:

Then fix the cause, not the symptom. In the example above, the fix was

separating the recorders so the assertion had two distinguishable sides β€” not

a sharper assertion on a fixture that couldn't tell them apart.

And re-run the mutation afterwards. A rewrite isn't finished until it goes red;

more than once I found that my "fixed" version still survived, for a second

reason I hadn't spotted.

That happened to the sample project for this post, too. One of the "fixed"

tests turned out to be partly hollow: the assertion was right, but the fixture

only exercised a single order line, so a bug that dropped every line after the

first went unnoticed. It was caught by running a mutation against examples

written specifically to demonstrate this failure mode, by someone actively

looking for it. Knowing the shape isn't the same as proving the test can

fail.

One detail worth making explicit, because it's easy to get wrong in both

directions.

A mutation must be a valid wrong implementation: it compiles, it runs,

it's just incorrect. If you change a function name to one that doesn't exist,

every test touching that code fails β€” the hollow ones included. The build

going red proves the line runs, not that any assertion checks it.

The same applies to red phases generally. A test that fails because the

function doesn't exist yet is a weaker signal than one that fails on an

assertion. Sometimes that's unavoidable early in a cycle; it's worth noting

when it happens and proving the test properly once the code is there.

In the sample project, the runner enforces both rules: a mutation has to

type-check, and only an assertion failure counts as a kill.

Most agents can work plan-first: the agent produces a plan and nothing changes

in the code until the plan is agreed. It's worth doing, and worth structuring

the plan around Red / Green / Refactor per behaviour, with every test named and

its assertions stated. That forces a commitment to what red looks like before

anything is written, which is most of the value of TDD arriving before the

first line of code.

It also moves the same question one stage earlier. A plan is an artifact too,

and it comes back with recognisable shapes: a decision left open with

"either… or", a placeholder where a test should be named, a checklist item

ticked before any work exists β€” sometimes against a project-specific rule the

agent defined for itself rather than looking up.

The useful move is the same one: verify the artifact, not the report of it.

MODE: Plan. VERIFY ONLY. Report, do not fix.
Read the plan file. Do not rely on your memory of the edits.
For each check, report PASS or FAIL and quote the exact lines
with line numbers as evidence.

Seventeen checks on one plan; eight failed on the first pass. Every fix was a

one-line edit once identified.

Every stage of this work produces an artifact the agent will also report on β€”

a test suite, a red commit, a plan, a checklist. Both checks in this post come

from the same instinct: read the artifact against something that could have

come out differently.

These come up often enough to be worth naming, and they're much easier to

catch when you're expecting them:

The check itself is the cheap part of that list. For comparison, a hook-based

guard runs a second model on every file write; the practitioner who set one up

measured ten minutes against four without it, and concluded the trade-off

wasn't worth it for that work. A mutation check is one broken line and one

test run, on changes you were reviewing anyway.

It's worth being clear about what this does and doesn't buy you.

A mutation check proves a test can fail for one change. It doesn't prove the

suite is complete, and it won't surface problems a fix causes elsewhere β€” a

change that's correct in isolation and wrong for the code around it still needs

a reviewer.

It also isn't worth doing everywhere. On small, visible work β€” a new field, a

validation rule, a copy change β€” reading the diff tells you everything the

mutation would. Reach for it where the code has indirection: injected

dependencies, factories, downstream systems, async paths, anything you can't

verify by eye. Those are the same places a hollow test is invisible in review,

which is not a coincidence.

If you take one thing from this, take the rules β€” they cost nothing once

written, and they prevent most of the shapes above. The check is for the

residue: the test that looks right, that a rule would not have caught, and that

few would question in review.

The loop gives you sequence: test first, one behaviour, red before green. The

rules deal with the shapes that recur. And then, occasionally, one more

question: could this have failed?

Sample project with all six test shapes and a runner that mutates one line at a time: https://github.com/thasnim-fluxone/tests-that-cannot-fail. npm run mutate shows every hollow test surviving and every fixed one caught.

── more in #ai-agents 4 stories Β· sorted by recency
── more on @ibm bob 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/tdd-with-coding-agen…] indexed:0 read:14min 2026-09-30 Β· β€”