{"slug": "tdd-with-coding-agents-write-the-rules-then-check-they-held", "title": "TDD With Coding Agents: Write the Rules, Then Check They Held", "summary": "A developer working with IBM Bob as a coding agent found that tests written by agents often pass regardless of whether the code works, and proposes a set of instruction-file rules plus a post-green check to catch non-discriminating tests. The work, reproducible via a sample GitHub project with six test shapes and a line-by-line mutation runner, argues that existing TDD tooling such as TDD Guard enforces test sequence but not whether a test can actually fail. The author notes the patterns appeared across languages, codebases and tools, including instruction mechanisms like AGENTS.md, .bob/rules, CLAUDE.md and .cursor/rules.", "body_md": "Test-driven development and coding agents fit together unusually well. A test\n\nis a precise specification the agent can check its own work against. The\n\nred-green-refactor cycle keeps each change small enough to review. And the\n\n\"run the tests, read the output, try again\" rhythm is exactly what these tools\n\nare good at.\n\nThe guidance on how to do it has matured quickly, and most of it is sound. But\n\nafter working this way across different codebases and languages, with the agent\n\nwriting both the tests and the code, I kept finding tests that passed whether\n\nthe code worked or not.\n\nMost of those were preventable. A handful of rules, written once into the\n\nagent's instructions, stop the common shapes before they are written. The rest\n\nneeded a check after green, because no rule anticipates every way a test can\n\nfail to discriminate.\n\nThis post is both: the rules worth writing down, and the thirty-second check\n\nfor what they miss.\n\nThe work behind this post used IBM Bob as the coding agent. Nothing here is\n\nspecific to it. Bob routes tasks across several frontier models rather than a\n\nsingle fixed one, so a single session may not even have used the same model\n\nthroughout — and the same patterns turned up regardless of language, codebase\n\nor tool.\n\n*Everything in this post is reproducible: [sample project on GitHub](https://github.com/thasnim-fluxone/tests-that-cannot-fail)\n— six test shapes and a runner that mutates one line at a time.*\n\nIt's worth tracing how the advice has developed, because each stage solved a\n\nreal problem.\n\n**Keep the tests in human hands.** The early guidance was that if a model\n\nwrites both the code and the tests, the same assumption ends up in both and\n\nthey agree with each other. So humans wrote the tests and the AI wrote the\n\ncode. That division worked, and it's still good advice when the spec matters\n\nmore than the speed.\n\n**Write the discipline down.** Agents can produce twenty tests in the time you\n\nwrite one, and most teams took that trade. The guidance shifted to enforcing\n\nthe cycle through whatever instruction mechanism the tool offers — an\n\n`AGENTS.md` in the project root now works across several of them, alongside\n\ntool-specific rules files like `.bob/rules` for IBM Bob, `CLAUDE.md` for Claude\n\nCode and `.cursor/rules` for Cursor — carrying the project's conventions and\n\nthe test-first rule, one behaviour per cycle, a human gate before\n\nimplementing, and, in the stronger versions, committing the red test so the\n\nfailing state lives in git history.\n\nKent Beck, who originated TDD, has published a system prompt along these lines: *always follow the cycle, write the simplest failing test first, implement the minimum\nneeded to pass.*\n\n**Enforce it mechanically.** The newest approach doesn't rely on the agent\n\nfollowing instructions at all. TDD Guard hooks every file write and blocks it\n\nunless the process was honoured: a failing test exists, you're on one test at\n\na time, you're working outside-in. Under the covers it spins up a second model\n\nas a judge on each edit, because \"did this follow TDD?\" is a fuzzy question\n\nthat is easier to ask a model than to encode in rules.\n\nEach stage is a genuine improvement, and the progression is the right one.\n\nPrompting alone tends to produce what one practitioner aptly called \"test\n\nfirst, not test-driven\" — all the tests, then all the code, in two large\n\nsteps. Instruction files make the cycle stick more often. Hooks make it stick unless\n\nthe judge misses something.\n\nEvery tool I looked at governs **sequence**: was the test written first, is it\n\none behaviour, did red precede green. That's the hard part to enforce, and they\n\nenforce it well.\n\nThe step I'm suggesting governs something different: **whether the resulting test can fail at all.**\n\nThese go in whatever instruction mechanism your tool offers, and they are the\n\nhigher-value half of this post. Written once, they make most of the shapes in\n\nthe next section much less likely, though no instruction is followed perfectly:\n\nThose rules prevent most of the shapes below. What follows is for the cases\n\nthey cannot anticipate.\n\nAfter green, before moving on:\n\nBreak one production line on purpose. Say in advance which test should fail.\n\nRun it. Confirm it fails **on an assertion**. Then restore the line.\n\nThat's mutation testing, done by hand, one line at a time, at review time\n\nrather than in CI. Note the direction: the code is correct, you introduce a\n\ndeliberate defect, and the test failing is the good outcome — it's the test\n\ndoing its job.\n\nA test that can't fail isn't a new problem, and it isn't specific to AI.\n\nMutation testing has existed for decades precisely because of it, and most TDD\n\nguides mention it — usually a line in a metrics section beside coverage\n\nthresholds, pointing at Stryker or mutmut. That framing makes it sound like a\n\nquarterly exercise. Used per change, it's smaller and more immediate: the one\n\nquestion that separates a test from a decoration.\n\nWhat's different with agents is that TDD normally guards against this, and the\n\nguard gets weaker. The red step is meant to be the proof — you watched the test\n\nfail, so it can fail. That holds when a person writes one test and sees it go\n\nred for the reason they expected. It holds less well when an agent produces the\n\ntest: the red often comes from the function not existing yet rather than from\n\nthe assertion, and the volume means few of them get inspected closely.\n\nBetter instructions narrow this a long way, and you should write them. But\n\ninstructions produce better tests; they do not produce proof that a given test\n\ncan fail. That distinction is the whole reason the check exists.\n\nSo the mutation asks what the red phase no longer reliably answers. A red test\n\nproves it failed *before the code existed* — when everything failed. This asks:\n\nnow that the code exists, would this test notice if it broke?\n\nMost of the time the answer is yes, it takes thirty seconds, and you move on.\n\nOccasionally it isn't, and those are the cases worth writing about.\n\nTwo agent-written tests for the same behaviour: the service forwards the\n\ncaller's correlation ID to a downstream supplier. Both green.\n\nBreak the one production line they exist to protect:\n\n```\n- new SupplierClient(supplierTransport, { forwardCorrelation: true });\n+ new SupplierClient(supplierTransport, { forwardCorrelation: false });\n```\n\nThe service no longer forwards the header. It still compiles, which matters —\n\na mutation that doesn't compile tells you nothing. Then run the tests:\n\n```\n✓ hollow: correlation ID reaches the supplier\n× fixed:  correlation ID reaches the supplier\n\nAssertionError: expected undefined to be 'abc'\n  41|   expect(supplierSide.callCount).toBe(1);\n  42|   expect(supplierSide.last()?.headers[HEADER]).toBe(\"abc\");\n```\n\nRestore the line, confirm green, move on. Total cost: about thirty seconds.\n\nThe reasoning behind it: two passing tests tell you the tests and the code\n\nagree. That's true when the code is right — and equally true when the test\n\ncan't tell the difference. From a green suite, those two situations look\n\nidentical.\n\nIntroducing a defect separates them. A test that genuinely checks the behaviour\n\nhas to notice, because its result depends on that behaviour. The hollow one\n\npassed with forwarding switched on and with it switched off, which means its\n\nresult never depended on forwarding at all.\n\nI've rebuilt each of these on a small fictional service in TypeScript so you\n\ncan run them (link at the end). Each has a **hollow** version that passes and\n\na **fixed** version that also passes. The difference only shows under mutation.\n\nNone of these come from carelessness. They're the kind of test a competent\n\ndeveloper writes and a reviewer approves.\n\nFour of the six are preventable by a rule. Two are not, and those are the ones\n\nthat justify the check. I've marked each.\n\n*Preventable by rule 5.*\n\nThe service should forward a caller's correlation ID to a downstream supplier.\n\nThe test records outbound requests and checks for the header:\n\n``` js\nconst shared = new Recorder();\nconst service = buildService(new Transport(shared));\nconst caller = new Transport(shared, (req) => service.handle(req)); // same recorder\n\ncaller.send({ path: \"/orders\", headers: { [HEADER]: \"abc\" }, body: ORDER });\n\nexpect(shared.requests.some((r) => r.headers[HEADER] === \"abc\")).toBe(true);\n```\n\nThe test's own client and the code under test share one recorder. The inbound\n\nrequest already carries the header, so the recorder always sees it — whether\n\nor not the service forwarded anything. Break the forwarding and the test stays\n\ngreen.\n\nThe assertion is fine. The fixture defeats it. The fix is separate recorders\n\nand an assertion on the side that matters:\n\n```\nexpect(supplierSide.callCount).toBe(1);\nexpect(supplierSide.last()?.headers[HEADER]).toBe(\"abc\");\n```\n\nThis is the one I'd least expect to find by reading. The fixture looks\n\ncorrect; only the mutation shows it isn't.\n\n*Preventable by rule 4.*\n\nThe test constructs the objects by hand, with the right settings, instead of\n\nusing the factory production uses. When the factory stops passing the setting,\n\nthe test doesn't notice — it never calls the factory.\n\nConvincing in review: real code, real requests, real assertions. Just not the\n\ncode that ships.\n\n*Preventable by rule 2.*\n\n``` js\nexpect(rec.requests.some((r) => HEADER in r.headers)).toBe(false);\n```\n\n`some()` over an empty array is `false`, so if the code path never runs, this\n\npasses. \"Nothing wrong happened\" and \"nothing happened\" look identical. One\n\nline fixes it: prove the calls were made first.\n\n*Preventable by rule 3.*\n\nAn audit trail always contains the inbound entry, so asserting it isn't empty\n\ncan't tell you whether the reservation was recorded. Assert the specific entry.\n\n*Not preventable by a rule. Nothing in the test looks wrong; the argument for removing the assertion is reasonable until a mutation answers it.*\n\nA log's sequence number must advance only after a write succeeds. After a\n\nfailed write, the test checks the store is empty — which it always is, because\n\nthe write is all-or-nothing.\n\nThere's a reasonable argument for dropping the second assertion, that the\n\ncounter hasn't moved: the store is empty, so what could be wrong? The mutation\n\nanswers it. Advance the counter before the write, and the store is still\n\nempty, but:\n\n```\nexpected 2 to deeply equal +0\n```\n\nThe counter has moved past entries that were never written. Only the\n\n\"redundant\" assertion sees it.\n\n*Not preventable by a rule. The test is correct; the behaviour is invisible from where it is looking, and you only find that out by breaking the code.*\n\nOnly the first request of an order should carry a parent ID. The integration\n\ntest checks what the downstream system recorded — but that system keeps the\n\nvalue from the first call and ignores it afterwards. Sending it every time\n\nchanges nothing observable, and every integration test passes either way.\n\nThe integration test isn't wrong. It's looking where the behaviour is\n\ninvisible. A unit test on the outbound requests can see it.\n\nThe first instinct is to strengthen the assertion. That's usually the wrong\n\nmove, and it cost me two rewrites before I stopped reaching for it. A surviving\n\nmutation means the test's result didn't depend on the behaviour — so the\n\nquestion is *why not*, and the answer is often somewhere other than the\n\nassertion.\n\nFour things to check, in order:\n\nThen fix the *cause*, not the symptom. In the example above, the fix was\n\nseparating the recorders so the assertion had two distinguishable sides — not\n\na sharper assertion on a fixture that couldn't tell them apart.\n\nAnd re-run the mutation afterwards. A rewrite isn't finished until it goes red;\n\nmore than once I found that my \"fixed\" version still survived, for a second\n\nreason I hadn't spotted.\n\nThat happened to the sample project for this post, too. One of the \"fixed\"\n\ntests turned out to be partly hollow: the assertion was right, but the fixture\n\nonly exercised a single order line, so a bug that dropped every line after the\n\nfirst went unnoticed. It was caught by running a mutation against examples\n\nwritten specifically to demonstrate this failure mode, by someone actively\n\nlooking for it. Knowing the shape isn't the same as proving the test can\n\nfail.\n\nOne detail worth making explicit, because it's easy to get wrong in both\n\ndirections.\n\nA mutation must be a **valid wrong implementation**: it compiles, it runs,\n\nit's just incorrect. If you change a function name to one that doesn't exist,\n\nevery test touching that code fails — the hollow ones included. The build\n\ngoing red proves the line runs, not that any assertion checks it.\n\nThe same applies to red phases generally. A test that fails because the\n\nfunction doesn't exist yet is a weaker signal than one that fails on an\n\nassertion. Sometimes that's unavoidable early in a cycle; it's worth noting\n\nwhen it happens and proving the test properly once the code is there.\n\nIn the sample project, the runner enforces both rules: a mutation has to\n\ntype-check, and only an assertion failure counts as a kill.\n\nMost agents can work plan-first: the agent produces a plan and nothing changes\n\nin the code until the plan is agreed. It's worth doing, and worth structuring\n\nthe plan around Red / Green / Refactor per behaviour, with every test named and\n\nits assertions stated. That forces a commitment to what red looks like before\n\nanything is written, which is most of the value of TDD arriving before the\n\nfirst line of code.\n\nIt also moves the same question one stage earlier. A plan is an artifact too,\n\nand it comes back with recognisable shapes: a decision left open with\n\n\"either… or\", a placeholder where a test should be named, a checklist item\n\nticked before any work exists — sometimes against a project-specific rule the\n\nagent defined for itself rather than looking up.\n\nThe useful move is the same one: verify the artifact, not the report of it.\n\n```\nMODE: Plan. VERIFY ONLY. Report, do not fix.\nRead the plan file. Do not rely on your memory of the edits.\nFor each check, report PASS or FAIL and quote the exact lines\nwith line numbers as evidence.\n```\n\nSeventeen checks on one plan; eight failed on the first pass. Every fix was a\n\none-line edit once identified.\n\nEvery stage of this work produces an artifact the agent will also report on —\n\na test suite, a red commit, a plan, a checklist. Both checks in this post come\n\nfrom the same instinct: read the artifact against something that could have\n\ncome out differently.\n\nThese come up often enough to be worth naming, and they're much easier to\n\ncatch when you're expecting them:\n\nThe check itself is the cheap part of that list. For comparison, a hook-based\n\nguard runs a second model on every file write; the practitioner who set one up\n\nmeasured ten minutes against four without it, and concluded the trade-off\n\nwasn't worth it for that work. A mutation check is one broken line and one\n\ntest run, on changes you were reviewing anyway.\n\nIt's worth being clear about what this does and doesn't buy you.\n\nA mutation check proves a test can fail for *one* change. It doesn't prove the\n\nsuite is complete, and it won't surface problems a fix causes elsewhere — a\n\nchange that's correct in isolation and wrong for the code around it still needs\n\na reviewer.\n\nIt also isn't worth doing everywhere. On small, visible work — a new field, a\n\nvalidation rule, a copy change — reading the diff tells you everything the\n\nmutation would. Reach for it where the code has indirection: injected\n\ndependencies, factories, downstream systems, async paths, anything you can't\n\nverify by eye. Those are the same places a hollow test is invisible in review,\n\nwhich is not a coincidence.\n\nIf you take one thing from this, take the rules — they cost nothing once\n\nwritten, and they prevent most of the shapes above. The check is for the\n\nresidue: the test that looks right, that a rule would not have caught, and that\n\nfew would question in review.\n\nThe loop gives you sequence: test first, one behaviour, red before green. The\n\nrules deal with the shapes that recur. And then, occasionally, one more\n\nquestion: *could this have failed?*\n\n*Sample project with all six test shapes and a runner that mutates one line at\na time: [https://github.com/thasnim-fluxone/tests-that-cannot-fail](https://github.com/thasnim-fluxone/tests-that-cannot-fail).\n`npm run mutate` shows every hollow test surviving and every fixed one caught.*", "url": "https://wpnews.pro/news/tdd-with-coding-agents-write-the-rules-then-check-they-held", "canonical_source": "https://dev.to/thasnimfluxone/tdd-with-coding-agents-write-the-rules-then-check-they-held-29ak", "published_at": "2026-09-30 11:08:24+00:00", "updated_at": "2026-09-30 11:18:35.054193+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "developer-tools", "large-language-models"], "entities": ["IBM Bob", "TDD Guard", "Kent Beck", "GitHub", "Claude Code", "Cursor"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/tdd-with-coding-agents-write-the-rules-then-check-they-held", "markdown": "https://wpnews.pro/news/tdd-with-coding-agents-write-the-rules-then-check-they-held.md", "text": "https://wpnews.pro/news/tdd-with-coding-agents-write-the-rules-then-check-they-held.txt", "jsonld": "https://wpnews.pro/news/tdd-with-coding-agents-write-the-rules-then-check-they-held.jsonld"}}