Stop letting your AI agents hallucinate test failures A developer built QA Arbiter, a tool that prevents AI agents from hallucinating test failures by enforcing structured diagnostic workflows. The tool uses a 'Decision Pivot' pattern to force agents to mathematically trace test failures, distinguishing between engine defects and test errors, thus avoiding agentic deadlocks in CI pipelines. I've seen enough CI pipelines die in an infinite loop of 'fix, retry, fail' to last a lifetime. You know the pattern. An agent-driven QA process runs a suite. A Vitest assertion fails. The LLM looks at the error log, reads the code, and makes an executive decision: 'The logic is flawed; let me refactor this function.' It pushes a change. The pipeline runs again. It fails again with the exact same error—or worse, a slightly different one that breaks something else. You''re not looking at a debugging problem anymore; you're looking at an agentic deadlock. The fundamental issue isn't that LLMs are 'stupid.' It's actually more subtle: they are too confident in their own unverified reasoning. When you tell an agent to 'check the logs,' it doesn't actually perform a step-by-step arithmetic trace of the logic. It performs a high-level semantic scan, matches patterns that look like common bugs, and then reports back with total certainty. If the test author wrote expected: 108 instead of expected: 100 , the agent won't tell you the test is wrong. It will 'fix' your business logic to match a broken assertion. This is how 'coverage theater' turns into actual production regressions. This is why I built QA Arbiter https://vinkius.com/mcp/qa-arbiter . In a multi-agent pipeline—where one agent writes tests QA and another fixes code Dev —the cost of a single wrong diagnosis is massive. If the QA agent identifies an ENGINE DEFIT when it's actually a TEST ERROR , the Dev agent wastes cycles refactoring perfectly functional code. The developer loses trust in the automation, and suddenly your 'autonomous pipeline' becomes just another source of noise in Slack. I call this 'Assertion Confusion.' It’s common in cases like tax calculations or time-zone handling the classic midnight crossover bug . An agent sees a failure, assumes there is a math error in the code, and implements a fix that actually changes the domain logic to satisfy an incorrect test expectation. Without a way to force the agent to prove its work, you are just automating chaos. QA Arbiter doesn't run your tests. It doesn't compute values. If it did, I would have just written a Python script and called it a day. Instead, QA Arbiter acts as a reasoning enforcer using the Decision Pivot pattern. When an agent encounters a failure, it is required to call diagnose test failure . This isn't an optional suggestion; the tool structure forces the agent into a deterministic workflow: receivedMatchesTrace and expectedMatchesTrace .The magic is in the verdict logic. There are no 'maybe's.' Received == Trace AND Expected = Trace , it is a Received = Trace AND Expected == Trace , it is an The tool validates logical consistency. If an agent tries to claim there is an ENGINE DEFECT but its own trace shows that the received value matches the calculation, the tool rejects the call. It forces the agent to re-examine the gap in its reasoning. If you are building production-grade MCP integrations via Vinkius, you know that 'instruction following' is a suggestion, but 'tool execution' is an obligation. We built QA Arbiter to prevent the specific type of regression that happens when developers trust agentic summaries too much. By using structured diagnostics, we move away from the LLM saying "I think this is the bug" and toward the LLM saying "Here is the mathematical trace that proves the engine failed at step 3." This eliminates the 'flaky test' tolerance that plagues modern CI. Instead of ignoring a failure because it 'might be flaky,' the agent can actually diagnose if the environment is polluted or if there's a race condition by tracing state changes through the tool call. If you are tired of your agents pushing 'fixes' that only serve to satisfy broken tests, you need to stop giving them instructions and start giving them constraints. Check out the full implementation here: https://vinkius.com/mcp/qa-arbiter https://vinkius.com/mcp/qa-arbiter MCPs are the music of AI Agents. We built the catalog. Discover Vinkius MCP Catalog.