{"slug": "red-green-remove-outside-in-tests-for-coding-agents", "title": "Red, Green, Remove: Outside-In Tests for Coding Agents", "summary": "The OpenClaw team deleted roughly 400,000 lines of its own tests with little change in code coverage, and published a test-audit skill that gates new tests with four questions and sweeps existing tests for junk. A separate outside-in-tests skill proposes a \"red, green, remove\" workflow in which developers keep only behavior tests through the real entry point, edge-case tables, contract checks, and regression tests, and delete the rest before landing. The author argues line coverage only maps untested code and recommends mutation testing to determine whether a test would actually fail if the code broke.", "body_md": "# Red, Green, Remove: Outside-In Tests for Coding Agents\n\ncoding agentstestingskills\n\nThe OpenClaw team reported deleting around 400,000 lines of their own tests without much change in code coverage. That number stuck with me. Models love writing a test for every tiny change, and they rarely clean up after themselves.\n\nOpenClaw published the [test-audit skill](https://github.com/openclaw/openclaw/blob/main/.agents/skills/test-audit/SKILL.md) behind that cleanup. It gates new tests with four questions and sweeps the existing pile for junk. It’s useful, and I borrowed its questions. But it reacts. The better move is to stop making the pile.\n\n## Test sediment\n\nTest-driven development works well for agents. Write a failing test, make it pass, move on. The trouble comes after green. Each step leaves a small test behind. I call the pile test sediment.\n\nSediment slows every build, and agents run the suite constantly. It breaks on refactors that keep behavior. Worst, it can lock in wrong behavior. A bug fix then means defending the old wrong answer first.\n\n## Describe behavior first\n\nMy approach starts outside the code. Describe the behavior in the user’s words. Grow tests red-green. Keep what states behavior, and remove the rest before landing. The goal is one set of outside-in tests that checks behavior and edge cases. No duplicates, and nothing tautological.\n\nI’ve found BDD with Cucumber powerful for the first step. Given a customer with a paid order, when they ask for a refund, then the order shows a pending refund. Cucumber’s example tables become edge-case tables. The shape matters more than the tool. Glue code between the plain-language steps and the system can become sediment too, so keep it thin or skip the tool. A brief that lists behaviors this way hands an agent its outside-in tests.\n\n## Test from the outside in\n\nTest each behavior where the user or caller meets it: the command line, the API or the screen. Care little about the layers underneath. They may change while the behavior holds. Test something inside directly only when it owns behavior worth stating on its own, such as a parser.\n\nGive each behavior one owning test. A bug gets one regression test where it shows. Don’t repeat it at every layer the bug crosses.\n\n## Red, green, remove\n\n1. Write the behavior as an outside-in test. Watch it fail for the right reason.\n2. When a step is hard, drop to unit tests. They are scaffolding for the work in progress.\n3. Make it pass.\n4. Before landing, keep a test only if it is one of the four kinds below. Delete the rest.\n\nThe four kinds that stay:\n\n- **Behavior tests** through the real entry point.\n- **Edge-case tables** : boundaries, empty input, bad input and limits in one table.\n- **Contract checks** on what others depend on: schemas, output formats, secrecy, releases.\n- **Regression tests** that failed on the code before the fix.\n\n## Coverage counts lines, not claims\n\nLine coverage tells you which code ran during the tests. It doesn’t tell you whether any test would notice that code going wrong. A tautological test runs the code and asserts what its setup already made true. Coverage goes up, and nothing is checked. So treat coverage as a map of what’s untested, never as proof of what’s tested. Mutation testing asks the better question: if this code broke, would a test fail?\n\n## What about deleting all your unit tests?\n\nA lot of people now talk about deleting their unit tests to speed things up. They’re mostly right, but deleting all of them overcorrects. An edge-case table for a parser is one of the best tests you have. Delete with evidence. Mutation testing shows which tests catch mistakes nothing else catches. Never delete the only test of a behavior. Write the outside-in test first, then delete.\n\n## The skill\n\nI wrote this up as a skill you can drop into your own agent setup: [outside-in-tests.md](https://gist.github.com/imaurer/ac31f596bcfd7f46afe1c7dceedcba21). It pairs with OpenClaw’s test-audit. Theirs cleans up the pile, and this one keeps it from forming. It also pairs with my earlier post on [coding agent proof spirals](https://www.imaurer.com/writing/coding-agent-proof-spirals/).\n\nWhen someone says “we deleted all our unit tests and nothing broke”, ask how they would know.\n\n## Also: introducing ThinkThen\n\nI recently introduced ThinkThen. It answers typed questions about text and returns `true`, `false`, a label or a number. A failed call never looks like an answer. Use it to gate a script, label records, or grade answers in an eval.\n\n```\nthinkthen decide 'Does the customer ask for money back?' < message.txt\n```\n\nThe command prints `true` and exits 0 for yes, 1 for no and 3 for not sure. A shell `if` can branch on it. It has ten functions, such as `decide`, `choose`, `tag` and `score`. They run across 25 surfaces: the command line, a Rust crate, language bindings from Python to COBOL, and SQL extensions for DuckDB, SQLite and PostgreSQL.\n\nRead the docs at [thinkthen.dev](https://thinkthen.dev). The code lives at [github.com/botassembly/thinkthen](https://github.com/botassembly/thinkthen).", "url": "https://wpnews.pro/news/red-green-remove-outside-in-tests-for-coding-agents", "canonical_source": "https://www.imaurer.com/writing/red-green-remove-outside-in-tests/", "published_at": "2026-10-05 14:39:53+00:00", "updated_at": "2026-10-05 14:50:00.546110+00:00", "lang": "en", "topics": ["ai-agents", "developer-tools", "ai-tools"], "entities": ["OpenClaw", "Cucumber"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/red-green-remove-outside-in-tests-for-coding-agents", "markdown": "https://wpnews.pro/news/red-green-remove-outside-in-tests-for-coding-agents.md", "text": "https://wpnews.pro/news/red-green-remove-outside-in-tests-for-coding-agents.txt", "jsonld": "https://wpnews.pro/news/red-green-remove-outside-in-tests-for-coding-agents.jsonld"}}