cd /news/ai-tools/coverage-theatre-your-ai-hit-90-cove… · home › topics › ai-tools › article
[ARTICLE · art-142383] src=dev.to ↗ pub= topic=ai-tools verified=true sentiment=· neutral

Coverage theatre: your AI hit 90% coverage and still shipped the bug

A developer published a free, self-contained lab demonstrating that AI-generated test suites can reach high code coverage while still passing against deliberately broken code. In the lab, the AI's own suite passes 3/3 against a pricing function seeded with four defects, while an independent verification suite fails 8/12 and catches all four, with a mutation gate confirming the weak suite never goes red on broken code. The author argues coverage measures which lines ran, not whether a test would fail if the code were wrong, and recommends mutation testing as the cheapest check.

by read2 min views1 publishedSep 30, 2026

Your AI assistant just generated tests until the coverage bar turned green. 92%. Ship it?

Here's the trap: coverage measures which lines ran, not whether a test would notice when they're wrong. AI made that gap free — you can generate 90% coverage in a minute, all of it happy-path, none of it discriminating. That's coverage theatre: it looks like safety, it measures activity, it proves almost nothing.

A line is "covered" the moment a test executes it. But executing a line and asserting the right thing about it are different claims. An AI-generated test that calls applyDiscount(order) and asserts typeof result === 'number' covers the function — and would stay green if the discount math were completely wrong.

Coverage answers "did this code run in a test?" The question that earns trust is "would a test fail if this code were wrong?" Those are not the same, and AI output routinely nails the first while skipping the second.

If a suite can't tell correct behavior from a real defect, its coverage number is decoration. The cheapest check is mutation: change > to >=, flip a boolean, swap two lines — then see if anything goes red. If the suite stays green on broken code, it isn't protecting you.

I put the smallest runnable version of this in a free lab (no install, Node 20+):

node run-lab.js all

It hands the AI's own suite a pricing function with four seeded defects, then runs an independent verification suite against the same broken code:

WEAK TESTS (the AI's suite) .... PASS 3/3
STRONG VERIFICATION ........... FAIL 8/12   <- catches all 4 defects
MUTATION GATE ................. PASS 4/4
LAB GATE: PASS

The AI's tests pass against code that is wrong. Verification is what fails — and that failure is the signal coverage never gave you.

Coverage isn't useless — low coverage is a real red flag. But high coverage is not evidence of anything on its own, and AI made it trivially easy to manufacture. Treat the green bar as a starting question, not an answer.

Free, self-contained lab: https://github.com/chernevnikolay86-wq/ai-test-verification-lab

I turned the full method into a book + runnable kit — the operating loop, governance, and labs across unit, web/API, instrument-cluster and CAN/UDS domains.

Honest caveat: the lab systems are simulators with declared defects — they prove the method catches faults, not that any product is certified. That's the point: don't trust things that merely look right.

── more in #ai-tools 4 stories · sorted by recency
── more on @github 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/coverage-theatre-you…] indexed:0 read:2min 2026-09-30 · —