Originally published on hexisteme notes.
I run a review step that sends the same question to two models from different vendors and reads back structured verdicts. On one batch of four rulings the two legs picked different answers on two of them.
A 1β1 split between two voters is not a tie you can break by counting. It is a coin flip.
Both splits resolved cleanly anyway, and neither resolution involved the picks. It involved what both legs had thrown away.
The panel was supposed to have three legs. The third β a CLI worker from a third vendor β had hit its usage cap that morning, with a reset date three days out. I recorded the gap instead of quietly shipping a two-leg result as if it were the designed one, and then had to actually work out what a two-leg split means, because the majority I would normally have reached for did not exist.
It is worth saying plainly, because the arithmetic is easy to skip: majority voting needs at least three independent voters. With two, "2β0" is agreement and "1β1" carries no information if the only thing you recorded is the pick. The fix is not a third leg. The fix is to record more than the pick.
Every option in the brief is numbered, and every answer has to come back in one shape:
N. <choice> β <2β3 lines of reasoning> β falsified if: <one line>
That last field looks like paperwork and does half the work in this post. A model asked to name the condition under which its own answer would be wrong will, often enough to matter, name a condition that is already true. You cannot read that if you never asked for it, and no amount of re-reading the choice will recover it.
The question was whether an abrupt change in a character's behaviour needed setup on the page. Four options went out:
Leg one picked A. Leg two picked C.
Tally: 1β1, deadlock. Overlap: both picks contain A. C is A+B. The only thing actually in dispute was B.
B then lost to a domain invariant rather than a vote. It required inventing a fact the project's canon did not have, and the project has an explicit rule against speculative canon β not a style preference, but the constraint that keeps a long series from contradicting itself twenty chapters later. An option that can only be executed by breaking a standing invariant isn't a candidate; it's a bug report about the brief.
Then the falsification field paid for itself. The leg that chose C had written its own falsifier as, in substance, "wrong if the coincidence is later revealed as a deliberate setup by the adults." That is the central motif of the work. The leg had handed over the condition that invalidated its own pick, in the same answer, unprompted.
Three independent reasons converged on A. None of them was the tally.
Executing A meant opening the manuscript at the clause to be strengthened. It pointed at the wrong place: the clause named one location, and the scene it referenced happens somewhere else entirely.
The flow reviewer had read eight chapters end to end and had not caught it. That axis reads who, when, and why; it does not check where. A panel that agrees is still only as wide as the axes you gave it, and one leg can raise an objection without being able to settle one.
Second ruling: whether a transgression at the climax leaves a visible price on the page. Options:
Leg one picked A. Leg two picked D. 1β1 again.
Both rejected B and C β and rejected them for the same structural reason. B and C each reframe the event as an incomplete repayment, and an earlier ruling in the same project had fixed the opposite invariant: handing the object back is a registration, not a repayment. That distinction is the engine the whole series runs on. Two models arriving independently at "these two options switch off the engine" is a far stronger signal than either one's preference between the survivors.
That left A and D, and ground truth cut it. D asked for one more beat of the character noticing. Grep the chapter: the beat is already there, twice.
Across those four rulings one leg picked option A every single time β 4 for 4. The other picked A twice, C once, D once.
That is disposition. How interventionist a model is by default is a real, stable property, and it rides directly on the choice. If you tally choices from a two-model panel, a good fraction of what you are measuring is which vendor's model is more eager to change things.
It does not ride the same way on the rejection. In those four rulings neither model ever picked an option the other had explicitly ruled out. The two unanimous rulings were unanimous on the reject side too: in one, both legs named the same bad option β make a required phrase "appear" by relocating it into an unrelated scene β and gave nearly the same reason. The string crosses; the institution behind the string does not.
Four rulings is an anecdote, not a study, and I'll take the correction if the next twenty go the other way. What makes me willing to act on it now is the mechanism rather than the count. A pick is one sample from a distribution that training shaped. A rejection with a cited reason is a claim that a specific constraint was checked and failed. Those are different kinds of statement, and counting them in the same unit is the actual error.
Same week, a different gate, two legs split on a diagnosis rather than a menu. The question was whether a batch of worldbuilding passages sat inert next to the main conflict. One leg read eight of eight sampled passages as unrelated to conflict. The other read five of eight as conflict-generating and named which conflict each one fed β this rule powers that theft, that rule powers that concealment.
There was no shared rejection to read, because there was no menu to reject from. The tiebreak was evidence: one leg cited passages, the other asserted a summary judgement.
Citations beat verbosity, and this is the tiebreak most worth naming out loud, because length is not evidence and the longer answer always feels more considered. The cited answer also changed the diagnosis rather than settling it: the disease was not inert exposition. It was re-introducing a rule that had already done work in an earlier chapter β a duplicate, not a filler. No vote-counting rule would have produced that, because the winning answer wasn't one of the options.
| Ruling | Leg A picked | Leg B picked | Both discarded | Still standing |
|---|---|---|---|---|
| Climax price | A | D | B Β· C | A Β· D |
Two human reviewers on a pull request. Two static analysers with overlapping rule sets. Two vendors answering the same architecture question. The move is the same: ask each one what they would rule out and why, not only what they would do.
Vetoes intersect more cleanly than recommendations, because a recommendation is partly a statement about the recommender and a veto with a reason attached is a claim about the artifact. One of those is auditable. And when you have exactly two reviewers, auditable is all you have β there is no majority to hide behind.
More notes at hexisteme.github.io/notes.