TDD With Coding Agents: Write the Rules, Then Check They Held A developer working with IBM Bob as a coding agent found that tests written by agents often pass regardless of whether the code works, and proposes a set of instruction-file rules plus a post-green check to catch non-discriminating tests. The work, reproducible via a sample GitHub project with six test shapes and a line-by-line mutation runner, argues that existing TDD tooling such as TDD Guard enforces test sequence but not whether a test can actually fail. The author notes the patterns appeared across languages, codebases and tools, including instruction mechanisms like AGENTS.md, .bob/rules, CLAUDE.md and .cursor/rules. Test-driven development and coding agents fit together unusually well. A test is a precise specification the agent can check its own work against. The red-green-refactor cycle keeps each change small enough to review. And the "run the tests, read the output, try again" rhythm is exactly what these tools are good at. The guidance on how to do it has matured quickly, and most of it is sound. But after working this way across different codebases and languages, with the agent writing both the tests and the code, I kept finding tests that passed whether the code worked or not. Most of those were preventable. A handful of rules, written once into the agent's instructions, stop the common shapes before they are written. The rest needed a check after green, because no rule anticipates every way a test can fail to discriminate. This post is both: the rules worth writing down, and the thirty-second check for what they miss. The work behind this post used IBM Bob as the coding agent. Nothing here is specific to it. Bob routes tasks across several frontier models rather than a single fixed one, so a single session may not even have used the same model throughout — and the same patterns turned up regardless of language, codebase or tool. Everything in this post is reproducible: sample project on GitHub https://github.com/thasnim-fluxone/tests-that-cannot-fail — six test shapes and a runner that mutates one line at a time. It's worth tracing how the advice has developed, because each stage solved a real problem. Keep the tests in human hands. The early guidance was that if a model writes both the code and the tests, the same assumption ends up in both and they agree with each other. So humans wrote the tests and the AI wrote the code. That division worked, and it's still good advice when the spec matters more than the speed. Write the discipline down. Agents can produce twenty tests in the time you write one, and most teams took that trade. The guidance shifted to enforcing the cycle through whatever instruction mechanism the tool offers — an AGENTS.md in the project root now works across several of them, alongside tool-specific rules files like .bob/rules for IBM Bob, CLAUDE.md for Claude Code and .cursor/rules for Cursor — carrying the project's conventions and the test-first rule, one behaviour per cycle, a human gate before implementing, and, in the stronger versions, committing the red test so the failing state lives in git history. Kent Beck, who originated TDD, has published a system prompt along these lines: always follow the cycle, write the simplest failing test first, implement the minimum needed to pass. Enforce it mechanically. The newest approach doesn't rely on the agent following instructions at all. TDD Guard hooks every file write and blocks it unless the process was honoured: a failing test exists, you're on one test at a time, you're working outside-in. Under the covers it spins up a second model as a judge on each edit, because "did this follow TDD?" is a fuzzy question that is easier to ask a model than to encode in rules. Each stage is a genuine improvement, and the progression is the right one. Prompting alone tends to produce what one practitioner aptly called "test first, not test-driven" — all the tests, then all the code, in two large steps. Instruction files make the cycle stick more often. Hooks make it stick unless the judge misses something. Every tool I looked at governs sequence : was the test written first, is it one behaviour, did red precede green. That's the hard part to enforce, and they enforce it well. The step I'm suggesting governs something different: whether the resulting test can fail at all. These go in whatever instruction mechanism your tool offers, and they are the higher-value half of this post. Written once, they make most of the shapes in the next section much less likely, though no instruction is followed perfectly: Those rules prevent most of the shapes below. What follows is for the cases they cannot anticipate. After green, before moving on: Break one production line on purpose. Say in advance which test should fail. Run it. Confirm it fails on an assertion . Then restore the line. That's mutation testing, done by hand, one line at a time, at review time rather than in CI. Note the direction: the code is correct, you introduce a deliberate defect, and the test failing is the good outcome — it's the test doing its job. A test that can't fail isn't a new problem, and it isn't specific to AI. Mutation testing has existed for decades precisely because of it, and most TDD guides mention it — usually a line in a metrics section beside coverage thresholds, pointing at Stryker or mutmut. That framing makes it sound like a quarterly exercise. Used per change, it's smaller and more immediate: the one question that separates a test from a decoration. What's different with agents is that TDD normally guards against this, and the guard gets weaker. The red step is meant to be the proof — you watched the test fail, so it can fail. That holds when a person writes one test and sees it go red for the reason they expected. It holds less well when an agent produces the test: the red often comes from the function not existing yet rather than from the assertion, and the volume means few of them get inspected closely. Better instructions narrow this a long way, and you should write them. But instructions produce better tests; they do not produce proof that a given test can fail. That distinction is the whole reason the check exists. So the mutation asks what the red phase no longer reliably answers. A red test proves it failed before the code existed — when everything failed. This asks: now that the code exists, would this test notice if it broke? Most of the time the answer is yes, it takes thirty seconds, and you move on. Occasionally it isn't, and those are the cases worth writing about. Two agent-written tests for the same behaviour: the service forwards the caller's correlation ID to a downstream supplier. Both green. Break the one production line they exist to protect: - new SupplierClient supplierTransport, { forwardCorrelation: true } ; + new SupplierClient supplierTransport, { forwardCorrelation: false } ; The service no longer forwards the header. It still compiles, which matters — a mutation that doesn't compile tells you nothing. Then run the tests: ✓ hollow: correlation ID reaches the supplier × fixed: correlation ID reaches the supplier AssertionError: expected undefined to be 'abc' 41| expect supplierSide.callCount .toBe 1 ; 42| expect supplierSide.last ?.headers HEADER .toBe "abc" ; Restore the line, confirm green, move on. Total cost: about thirty seconds. The reasoning behind it: two passing tests tell you the tests and the code agree. That's true when the code is right — and equally true when the test can't tell the difference. From a green suite, those two situations look identical. Introducing a defect separates them. A test that genuinely checks the behaviour has to notice, because its result depends on that behaviour. The hollow one passed with forwarding switched on and with it switched off, which means its result never depended on forwarding at all. I've rebuilt each of these on a small fictional service in TypeScript so you can run them link at the end . Each has a hollow version that passes and a fixed version that also passes. The difference only shows under mutation. None of these come from carelessness. They're the kind of test a competent developer writes and a reviewer approves. Four of the six are preventable by a rule. Two are not, and those are the ones that justify the check. I've marked each. Preventable by rule 5. The service should forward a caller's correlation ID to a downstream supplier. The test records outbound requests and checks for the header: js const shared = new Recorder ; const service = buildService new Transport shared ; const caller = new Transport shared, req = service.handle req ; // same recorder caller.send { path: "/orders", headers: { HEADER : "abc" }, body: ORDER } ; expect shared.requests.some r = r.headers HEADER === "abc" .toBe true ; The test's own client and the code under test share one recorder. The inbound request already carries the header, so the recorder always sees it — whether or not the service forwarded anything. Break the forwarding and the test stays green. The assertion is fine. The fixture defeats it. The fix is separate recorders and an assertion on the side that matters: expect supplierSide.callCount .toBe 1 ; expect supplierSide.last ?.headers HEADER .toBe "abc" ; This is the one I'd least expect to find by reading. The fixture looks correct; only the mutation shows it isn't. Preventable by rule 4. The test constructs the objects by hand, with the right settings, instead of using the factory production uses. When the factory stops passing the setting, the test doesn't notice — it never calls the factory. Convincing in review: real code, real requests, real assertions. Just not the code that ships. Preventable by rule 2. js expect rec.requests.some r = HEADER in r.headers .toBe false ; some over an empty array is false , so if the code path never runs, this passes. "Nothing wrong happened" and "nothing happened" look identical. One line fixes it: prove the calls were made first. Preventable by rule 3. An audit trail always contains the inbound entry, so asserting it isn't empty can't tell you whether the reservation was recorded. Assert the specific entry. Not preventable by a rule. Nothing in the test looks wrong; the argument for removing the assertion is reasonable until a mutation answers it. A log's sequence number must advance only after a write succeeds. After a failed write, the test checks the store is empty — which it always is, because the write is all-or-nothing. There's a reasonable argument for dropping the second assertion, that the counter hasn't moved: the store is empty, so what could be wrong? The mutation answers it. Advance the counter before the write, and the store is still empty, but: expected 2 to deeply equal +0 The counter has moved past entries that were never written. Only the "redundant" assertion sees it. Not preventable by a rule. The test is correct; the behaviour is invisible from where it is looking, and you only find that out by breaking the code. Only the first request of an order should carry a parent ID. The integration test checks what the downstream system recorded — but that system keeps the value from the first call and ignores it afterwards. Sending it every time changes nothing observable, and every integration test passes either way. The integration test isn't wrong. It's looking where the behaviour is invisible. A unit test on the outbound requests can see it. The first instinct is to strengthen the assertion. That's usually the wrong move, and it cost me two rewrites before I stopped reaching for it. A surviving mutation means the test's result didn't depend on the behaviour — so the question is why not , and the answer is often somewhere other than the assertion. Four things to check, in order: Then fix the cause , not the symptom. In the example above, the fix was separating the recorders so the assertion had two distinguishable sides — not a sharper assertion on a fixture that couldn't tell them apart. And re-run the mutation afterwards. A rewrite isn't finished until it goes red; more than once I found that my "fixed" version still survived, for a second reason I hadn't spotted. That happened to the sample project for this post, too. One of the "fixed" tests turned out to be partly hollow: the assertion was right, but the fixture only exercised a single order line, so a bug that dropped every line after the first went unnoticed. It was caught by running a mutation against examples written specifically to demonstrate this failure mode, by someone actively looking for it. Knowing the shape isn't the same as proving the test can fail. One detail worth making explicit, because it's easy to get wrong in both directions. A mutation must be a valid wrong implementation : it compiles, it runs, it's just incorrect. If you change a function name to one that doesn't exist, every test touching that code fails — the hollow ones included. The build going red proves the line runs, not that any assertion checks it. The same applies to red phases generally. A test that fails because the function doesn't exist yet is a weaker signal than one that fails on an assertion. Sometimes that's unavoidable early in a cycle; it's worth noting when it happens and proving the test properly once the code is there. In the sample project, the runner enforces both rules: a mutation has to type-check, and only an assertion failure counts as a kill. Most agents can work plan-first: the agent produces a plan and nothing changes in the code until the plan is agreed. It's worth doing, and worth structuring the plan around Red / Green / Refactor per behaviour, with every test named and its assertions stated. That forces a commitment to what red looks like before anything is written, which is most of the value of TDD arriving before the first line of code. It also moves the same question one stage earlier. A plan is an artifact too, and it comes back with recognisable shapes: a decision left open with "either… or", a placeholder where a test should be named, a checklist item ticked before any work exists — sometimes against a project-specific rule the agent defined for itself rather than looking up. The useful move is the same one: verify the artifact, not the report of it. MODE: Plan. VERIFY ONLY. Report, do not fix. Read the plan file. Do not rely on your memory of the edits. For each check, report PASS or FAIL and quote the exact lines with line numbers as evidence. Seventeen checks on one plan; eight failed on the first pass. Every fix was a one-line edit once identified. Every stage of this work produces an artifact the agent will also report on — a test suite, a red commit, a plan, a checklist. Both checks in this post come from the same instinct: read the artifact against something that could have come out differently. These come up often enough to be worth naming, and they're much easier to catch when you're expecting them: The check itself is the cheap part of that list. For comparison, a hook-based guard runs a second model on every file write; the practitioner who set one up measured ten minutes against four without it, and concluded the trade-off wasn't worth it for that work. A mutation check is one broken line and one test run, on changes you were reviewing anyway. It's worth being clear about what this does and doesn't buy you. A mutation check proves a test can fail for one change. It doesn't prove the suite is complete, and it won't surface problems a fix causes elsewhere — a change that's correct in isolation and wrong for the code around it still needs a reviewer. It also isn't worth doing everywhere. On small, visible work — a new field, a validation rule, a copy change — reading the diff tells you everything the mutation would. Reach for it where the code has indirection: injected dependencies, factories, downstream systems, async paths, anything you can't verify by eye. Those are the same places a hollow test is invisible in review, which is not a coincidence. If you take one thing from this, take the rules — they cost nothing once written, and they prevent most of the shapes above. The check is for the residue: the test that looks right, that a rule would not have caught, and that few would question in review. The loop gives you sequence: test first, one behaviour, red before green. The rules deal with the shapes that recur. And then, occasionally, one more question: could this have failed? Sample project with all six test shapes and a runner that mutates one line at a time: https://github.com/thasnim-fluxone/tests-that-cannot-fail https://github.com/thasnim-fluxone/tests-that-cannot-fail . npm run mutate shows every hollow test surviving and every fixed one caught.