coding agentstestingskills
The OpenClaw team reported deleting around 400,000 lines of their own tests without much change in code coverage. That number stuck with me. Models love writing a test for every tiny change, and they rarely clean up after themselves.
OpenClaw published the test-audit skill behind that cleanup. It gates new tests with four questions and sweeps the existing pile for junk. It’s useful, and I borrowed its questions. But it reacts. The better move is to stop making the pile.
Test sediment #
Test-driven development works well for agents. Write a failing test, make it pass, move on. The trouble comes after green. Each step leaves a small test behind. I call the pile test sediment.
Sediment slows every build, and agents run the suite constantly. It breaks on refactors that keep behavior. Worst, it can lock in wrong behavior. A bug fix then means defending the old wrong answer first.
Describe behavior first #
My approach starts outside the code. Describe the behavior in the user’s words. Grow tests red-green. Keep what states behavior, and remove the rest before landing. The goal is one set of outside-in tests that checks behavior and edge cases. No duplicates, and nothing tautological.
I’ve found BDD with Cucumber powerful for the first step. Given a customer with a paid order, when they ask for a refund, then the order shows a pending refund. Cucumber’s example tables become edge-case tables. The shape matters more than the tool. Glue code between the plain-language steps and the system can become sediment too, so keep it thin or skip the tool. A brief that lists behaviors this way hands an agent its outside-in tests.
Test from the outside in #
Test each behavior where the user or caller meets it: the command line, the API or the screen. Care little about the layers underneath. They may change while the behavior holds. Test something inside directly only when it owns behavior worth stating on its own, such as a parser.
Give each behavior one owning test. A bug gets one regression test where it shows. Don’t repeat it at every layer the bug crosses.
Red, green, remove #
- Write the behavior as an outside-in test. Watch it fail for the right reason.
- When a step is hard, drop to unit tests. They are scaffolding for the work in progress.
- Make it pass.
- Before landing, keep a test only if it is one of the four kinds below. Delete the rest.
The four kinds that stay:
- Behavior tests through the real entry point.
- Edge-case tables : boundaries, empty input, bad input and limits in one table.
- Contract checks on what others depend on: schemas, output formats, secrecy, releases.
- Regression tests that failed on the code before the fix.
Coverage counts lines, not claims #
Line coverage tells you which code ran during the tests. It doesn’t tell you whether any test would notice that code going wrong. A tautological test runs the code and asserts what its setup already made true. Coverage goes up, and nothing is checked. So treat coverage as a map of what’s untested, never as proof of what’s tested. Mutation testing asks the better question: if this code broke, would a test fail?
What about deleting all your unit tests? #
A lot of people now talk about deleting their unit tests to speed things up. They’re mostly right, but deleting all of them overcorrects. An edge-case table for a parser is one of the best tests you have. Delete with evidence. Mutation testing shows which tests catch mistakes nothing else catches. Never delete the only test of a behavior. Write the outside-in test first, then delete.
The skill #
I wrote this up as a skill you can drop into your own agent setup: outside-in-tests.md. It pairs with OpenClaw’s test-audit. Theirs cleans up the pile, and this one keeps it from forming. It also pairs with my earlier post on coding agent proof spirals.
When someone says “we deleted all our unit tests and nothing broke”, ask how they would know.
Also: introducing ThinkThen #
I recently introduced ThinkThen. It answers typed questions about text and returns true, false, a label or a number. A failed call never looks like an answer. Use it to gate a script, label records, or grade answers in an eval.
thinkthen decide 'Does the customer ask for money back?' < message.txt
The command prints true and exits 0 for yes, 1 for no and 3 for not sure. A shell if can branch on it. It has ten functions, such as decide, choose, tag and score. They run across 25 surfaces: the command line, a Rust crate, language bindings from Python to COBOL, and SQL extensions for DuckDB, SQLite and PostgreSQL.
Read the docs at thinkthen.dev. The code lives at github.com/botassembly/thinkthen.