cd /news/ai-agents/my-tests-still-passed-after-the-audi… · home › topics › ai-agents › article
[ARTICLE · art-142206] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

My Tests Still Passed After the Auditor Deleted the Division

A developer running a daily autonomous Claude Code agent on a single Windows PC reported that a separate auditor agent's mutation testing caught two changes whose tests passed even after the auditor deliberately broke the code. In one case, deleting a division that converted a 7-day visitor total into a daily average left all tests green because the test inputs (8 and 140) fell on the same side of the 20-visitor threshold either way; the fix added a total of 35 plus boundary cases at 19.9 and exactly 20 per day. A second function meant to reuse an existing overlap rule survived mutations that dropped the blog-topic field and shortened the 30-day window to 7, so the fake module was changed to record its call arguments and assert both fields and the period read from the source system's own constant.

by read3 min views4 publishedSep 30, 2026

A field note from the autonomous Claude Code agent I run every day on one Windows PC. The numbers come from its own ledgers, not from memory.

My agent never grades its own work. A separate auditor agent reads the diff, runs the tests, and returns PASS or FAIL. One of its checks is a small mutation test: it changes a line of the new code on purpose and runs the tests again. If they still pass, the tests weren't checking that line.

On one day it failed two of my changes that way, in two different parts of the system.

The agent was changing the advice an experiment report prints. The rule was: if a product page gets fewer than 20 visitors a day, don't touch the page yet, bring more people first. The number of visitors came from a 7-day total, so the new code divided it by the number of days.

The tests covered both branches. One used a 7-day total of 8, the other a total of 140.

Now delete the division and compare the raw totals with 20. 8 is still under. 140 is still over. Every test passes. The auditor deleted it, the tests stayed green, and it returned FAIL.

My own mutation checks had only flipped the comparison. The division was new too, and none of my mutations touched it.

The fix was one input where the two versions disagree: a total of 35. That is 5 a day (under 20) but more than 20 as a raw total. Plus a boundary pair: a total that rounds to 19.9 a day must stay under, exactly 20 must not.

Later that day, the agent wrote a function that borrows an overlap check from another system I run, one that publishes blog posts. The whole point of the function was "use exactly the same rule": compare both the title and the blog topic, over the last 30 days.

The test used a fake version of the other system and checked that overlapping titles were caught. That was all.

The auditor made two changes: pass only the title, and change 30 days to 7. Both survived. With the real ledger, the title-only version marked the next day's post as "no overlap" even though its blog topic overlapped. That was the exact failure the function was written to prevent.

The fix: the fake module now records what it was called with. The test asserts both fields were passed and that the period matches the other system's own constant, read directly from it, not copied.

Where this comes from. Every post here comes from one setup I run daily: a CLAUDE.md, memory files the agent reads before it touches anything, and a separate auditor agent that returns PASS or FAIL. The first 3 chapters of the book that walks through it are free as a PDF: https://dbsoul.gumroad.com/l/autonomous-ai-agents-claude-code-free-sample

The full edition is 11 chapters plus 4 ready-to-use templates (CLAUDE.md starter, memory files, auditor checklist, measurement guide) and a hands-on section for every chapter, $19 as a PDF: https://dbsoul.gumroad.com/l/autonomous-ai-agents-claude-code

Questions about the setup are welcome in the comments — I'll answer with what actually happened, not theory.

── more in #ai-agents 4 stories · sorted by recency
── more on @claude code 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/my-tests-still-passe…] indexed:0 read:3min 2026-09-30 · —