# Your coding agent tells you all tests pass. Sometimes that's not true.

> Source: <https://dev.to/rishi_g_25/your-coding-agent-tells-you-all-tests-pass-sometimes-thats-not-true-i32>
> Published: 2026-09-28 00:09:01+00:00

Coding agents are confident narrators. When Claude Code, Cursor, or similar tools finish a task, they give you a summary: what changed, what they ran, and whether it worked. The problem is that summary comes from the agent itself. Agents are often wrong about their own work, not maliciously, just optimistically.

Independent research on this ("From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents") found that across separate benchmarks, 44-76% of agent task failures involved the agent confidently reporting success anyway. Not edge cases but a substantial share of the time something goes wrong, the agent's own account tells you it did not.

Two patterns worth thinking about:

The subagent problem: Modern coding agents spin up their own helper agents to handle sub-tasks. Those subagents write their own transcripts, in their own files, which the main conversation never surfaces. If a subagent's test run fails, and the parent agent never actually reads that failure, the parent's closing summary can say all tests pass and actually believes it since it did not see the failure at all.

The quieter lie: This is the one that actually worries me most. It's not always a case of the agent failing to notice a problem but sometimes it doesn't fix the failing test. It adds a new, passing test right next to it, and reports the suite as green. Exit code says success. Nothing you would catch by just checking whether tests ran. You would need to know which specific test passed, not just that something did.

Why this matters more in unattended contexts. If you are reviewing every diff line by line, you might catch this. If your agent is running in CI, a scheduled job, or any pipeline where nobody is watching in real time, the agent's summary is the only account you get. There is no one there to notice something is off.

What I built to deal with this: a small, open-source tool called Rashomon. It hooks into Claude Code's tool-call lifecycle and keeps its own independent record of what actually ran, separate from whatever the agent's closing summary claims. When they disagree, it tells you. When they do not, it stays quiet. Most turns produce nothing at all, because most of the time, everything's fine, and a tool that is noisy on every run just trains you to ignore it.

It's free, Apache 2.0, local-only (no telemetry, no account, no content stored, and only identifiers and shapes). Currently Claude Code-specific, with broader support planned.

Repo: [https://github.com/altrace-dev-role/rashomon](https://github.com/altrace-dev-role/rashomon) 

I also am interested in hearing whether others have run into the confident but wrong problem and what you are doing about it.
