3-LLM Cross-Validation: A Consensus Mechanism A developer writing under the handle detective-noir described a 3-LLM cross-validation mechanism in the ATS-008 architecture that accepts a verdict only when three isolated models reach a 2/3 consensus, using AST-based semantic equivalence checking rather than word matching to judge agreement. The system weights each validator by track record and logs approvals, rejections and dissenting votes to a behavior log, tolerating exactly one compromised or hallucinating model. The stated goal is to ensure no single model can convict on its own in governance and audit pipelines. The most counterintuitive audit principle in the ATS-008 architecture is this: the more fluent a model's reasoning chain, the more you should suspect it was fabricated. A single LLM's verdict is, at best, an unverified witness statement . It can tell a perfectly coherent story about whether a piece of code is safe — and still be completely wrong. That is the confirmation-bias failure mode that single-model pipelines inherit silently. This post is part of the "Detective Reasoning + Behavior Logs + Technical Puzzles" series. We'll tear down the mechanism this architecture uses instead: 3-LLM cross-validation — three independent models, isolated from each other, whose conclusions are accepted only when they reach consensus. Every verification starts with create proposal . The question under test is registered as a proposal — who participates, whether a reference answer exists. From this moment on, every judgment is recorded. This is the detective's case file. self.consensus.create proposal proposal id=proposal id, content={"agent count": len outputs , "has reference": ...}, proposer id="cross validator", Three models answer the same question in isolation. The isolation matters: they don't know the others exist, so they can't collude or contaminate each other. If a reference answer exists e.g., a human-labeled security verdict , it becomes the baseline; otherwise the first model's output is the baseline for comparison. This is the most misunderstood part. The system does not count two models as agreeing because they used the same words — that would just be repetition. It runs semantic equivalence checking : both outputs are normalized AST normalization and structurally compared to decide whether they express the same conclusion. Model A says "this function has an integer overflow risk"; Model B says "a buffer overrun may occur here." Different words, same meaning → consistent. Listen to the facts, not the phrasing. equiv result = self.equivalence.check equivalence comparison base, output, threshold=similarity threshold Each verifier is not a single equal vote. The engine maintains per-validator weights — models with a stronger track record weigh more. Approvals and rejections are both recorded, each carrying its semantic similarity as a justification, written into the behavior log. The final gate is 2/3 : Consensus is only declared when strength exceeds 2/3 if consensus strength <= 2 / 3: do not flag Byzantine behavior Why 2/3? With three nodes, the system tolerates exactly one "traitor" — a compromised model, a down node, or one in full hallucination mode. As long as the other two are honest and agree, consensus holds. This is the simplest form of Byzantine fault tolerance: tolerate one bad witness, never trust two. If consensus is not reached, the proposal is rejected — and the rejection itself is appended to consensus history . The dissenting vote, the failure, the weight adjustment: all logged. In auditing, the question isn't "who was right" but "why did we rule this way." The point of this mechanism is not to pick the smartest model. It is to make sure no single model ever has the authority to convict on its own. In governance and audit pipelines, that matters more than raw accuracy — because auditing isn't about being usually right . It's about never allowing a single point of false testimony. Written by detective-noir — detective reasoning, behavior logs, and technical puzzles: the audit philosophy of the ATS-008 architecture.