Karpathy's llm-council showed a lot of people something quietly important: the
most useful part of asking five models a question is not reading five answers. It
is watching them review each other. His weekend project runs three stages — every
model answers, every model ranks the anonymised answers of the others, a chair
model synthesises — and people keep finding that the ranking stage surfaces things
the first stage never said.
We have been building in the same direction (AI Group Call — disclosure: our product; it runs each seat on a model you pick from the major
labs), and the pattern holds up in a different format: a live voice call where each participant
hears the whole conversation before its turn. Three things change when models stop
answering in isolation:
Positions move. In independent answers, nothing ever changes its mind. In a
shared conversation you see it happen: model B starts certain, hears model A's
constraint it hadn't considered, and walks back its own suggestion. That retraction
is the highest-signal moment in the whole session — it tells you which consideration
actually mattered, something five parallel chats never show you because they never
had to disagree.
Errors get caught by the room, not by you. One model hallucinating an API flag
tends to get corrected by another seat, roughly the way a colleague says "that
wasn't in the docs last I checked." It is not a substitute for verification, but it
beats you being the only reviewer of five confident paragraphs.
The format forces brevity. Spoken turns are short. Counter-intuitively, that is
a feature for decisions: you get "no, because X" instead of four thousand words of
"it depends." For long-form analysis, written councils still win — a voice room is
a debate club, not a research library.
Running the room is a skill. What we have seen work across hundreds of sessions:
Not a benchmark — confident is not correct, and a model agreeing with you is not
evidence. Not a code reviewer; spoken turns are too short for long proofs. Not a
legal or financial board. The niche where a live multi-model room genuinely beats
both single-model chat and multi-tab comparison is decisions with trade-offs: architecture calls, pricing questions, "ship or polish" arguments, rehearsing a
pitch against investors who interrupt.
Whatever tool you reach for, the takeaway is architectural, not product-specific:
the unit of AI leverage is moving from the answer to the argument between answers.
Build your workflow around getting that argument — cross-review, ranked dissent,
or just a room where the models can hear each other — and the model you pick
matters a lot less than it feels like it should.