When Two AI Reviewers Disagree, Read What They Both Rejected A developer's review system that sends the same question to two AI models from different vendors found that when the models disagreed, the most reliable signal came not from their picks but from what both rejected. In two deadlocked rulings, the developer resolved the tie by examining the falsification conditions the models provided and by applying domain invariants, rather than relying on majority vote. The developer emphasizes that majority voting requires at least three independent voters and that recording more than just the pick is essential. Originally published on hexisteme notes. I run a review step that sends the same question to two models from different vendors and reads back structured verdicts. On one batch of four rulings the two legs picked different answers on two of them. A 1–1 split between two voters is not a tie you can break by counting. It is a coin flip. Both splits resolved cleanly anyway, and neither resolution involved the picks. It involved what both legs had thrown away. The panel was supposed to have three legs. The third — a CLI worker from a third vendor — had hit its usage cap that morning, with a reset date three days out. I recorded the gap instead of quietly shipping a two-leg result as if it were the designed one, and then had to actually work out what a two-leg split means, because the majority I would normally have reached for did not exist. It is worth saying plainly, because the arithmetic is easy to skip: majority voting needs at least three independent voters. With two, "2–0" is agreement and "1–1" carries no information if the only thing you recorded is the pick . The fix is not a third leg. The fix is to record more than the pick. Every option in the brief is numbered, and every answer has to come back in one shape: N.