cd /news/ai-agents/red-green-remove-outside-in-tests-fo… · home › topics › ai-agents › article
[ARTICLE · art-145441] src=imaurer.com ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Red, Green, Remove: Outside-In Tests for Coding Agents

The OpenClaw team deleted roughly 400,000 lines of its own tests with little change in code coverage, and published a test-audit skill that gates new tests with four questions and sweeps existing tests for junk. A separate outside-in-tests skill proposes a "red, green, remove" workflow in which developers keep only behavior tests through the real entry point, edge-case tables, contract checks, and regression tests, and delete the rest before landing. The author argues line coverage only maps untested code and recommends mutation testing to determine whether a test would actually fail if the code broke.

read4 min views1 publishedOct 5, 2026
Red, Green, Remove: Outside-In Tests for Coding Agents
Image: source

coding agentstestingskills

The OpenClaw team reported deleting around 400,000 lines of their own tests without much change in code coverage. That number stuck with me. Models love writing a test for every tiny change, and they rarely clean up after themselves.

OpenClaw published the test-audit skill behind that cleanup. It gates new tests with four questions and sweeps the existing pile for junk. It’s useful, and I borrowed its questions. But it reacts. The better move is to stop making the pile.

Test sediment #

Test-driven development works well for agents. Write a failing test, make it pass, move on. The trouble comes after green. Each step leaves a small test behind. I call the pile test sediment.

Sediment slows every build, and agents run the suite constantly. It breaks on refactors that keep behavior. Worst, it can lock in wrong behavior. A bug fix then means defending the old wrong answer first.

Describe behavior first #

My approach starts outside the code. Describe the behavior in the user’s words. Grow tests red-green. Keep what states behavior, and remove the rest before landing. The goal is one set of outside-in tests that checks behavior and edge cases. No duplicates, and nothing tautological.

I’ve found BDD with Cucumber powerful for the first step. Given a customer with a paid order, when they ask for a refund, then the order shows a pending refund. Cucumber’s example tables become edge-case tables. The shape matters more than the tool. Glue code between the plain-language steps and the system can become sediment too, so keep it thin or skip the tool. A brief that lists behaviors this way hands an agent its outside-in tests.

Test from the outside in #

Test each behavior where the user or caller meets it: the command line, the API or the screen. Care little about the layers underneath. They may change while the behavior holds. Test something inside directly only when it owns behavior worth stating on its own, such as a parser.

Give each behavior one owning test. A bug gets one regression test where it shows. Don’t repeat it at every layer the bug crosses.

Red, green, remove #

  1. Write the behavior as an outside-in test. Watch it fail for the right reason.
  2. When a step is hard, drop to unit tests. They are scaffolding for the work in progress.
  3. Make it pass.
  4. Before landing, keep a test only if it is one of the four kinds below. Delete the rest.

The four kinds that stay:

  • Behavior tests through the real entry point.
  • Edge-case tables : boundaries, empty input, bad input and limits in one table.
  • Contract checks on what others depend on: schemas, output formats, secrecy, releases.
  • Regression tests that failed on the code before the fix.

Coverage counts lines, not claims #

Line coverage tells you which code ran during the tests. It doesn’t tell you whether any test would notice that code going wrong. A tautological test runs the code and asserts what its setup already made true. Coverage goes up, and nothing is checked. So treat coverage as a map of what’s untested, never as proof of what’s tested. Mutation testing asks the better question: if this code broke, would a test fail?

What about deleting all your unit tests? #

A lot of people now talk about deleting their unit tests to speed things up. They’re mostly right, but deleting all of them overcorrects. An edge-case table for a parser is one of the best tests you have. Delete with evidence. Mutation testing shows which tests catch mistakes nothing else catches. Never delete the only test of a behavior. Write the outside-in test first, then delete.

The skill #

I wrote this up as a skill you can drop into your own agent setup: outside-in-tests.md. It pairs with OpenClaw’s test-audit. Theirs cleans up the pile, and this one keeps it from forming. It also pairs with my earlier post on coding agent proof spirals.

When someone says “we deleted all our unit tests and nothing broke”, ask how they would know.

Also: introducing ThinkThen #

I recently introduced ThinkThen. It answers typed questions about text and returns true, false, a label or a number. A failed call never looks like an answer. Use it to gate a script, label records, or grade answers in an eval.

thinkthen decide 'Does the customer ask for money back?' < message.txt

The command prints true and exits 0 for yes, 1 for no and 3 for not sure. A shell if can branch on it. It has ten functions, such as decide, choose, tag and score. They run across 25 surfaces: the command line, a Rust crate, language bindings from Python to COBOL, and SQL extensions for DuckDB, SQLite and PostgreSQL.

Read the docs at thinkthen.dev. The code lives at github.com/botassembly/thinkthen.

── more in #ai-agents 4 stories · sorted by recency
── more on @openclaw 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/red-green-remove-out…] indexed:0 read:4min 2026-10-05 · —