What actually changes when AI models talk to each other before answering you A developer behind AI Group Call, a voice product that runs each participant on a user-selected model, reports that letting multiple models hear each other before answering changes outcomes in three ways: positions shift as models retract claims after hearing constraints from peers, hallucinations get corrected by other seats rather than the user, and spoken turns force brevity. The account builds on Andrej Karpathy's llm-council project, whose three-stage answer-rank-synthesize design showed that the ranking stage surfaces considerations the initial answers never raised. The author argues the unit of AI leverage is shifting from the answer to the argument between answers. Karpathy's llm-council showed a lot of people something quietly important: the most useful part of asking five models a question is not reading five answers. It is watching them review each other. His weekend project runs three stages — every model answers, every model ranks the anonymised answers of the others, a chair model synthesises — and people keep finding that the ranking stage surfaces things the first stage never said. We have been building in the same direction AI Group Call https://aigroupcall.app/llm-council-voice/ — disclosure: our product; it runs each seat on a model you pick from the major labs , and the pattern holds up in a different format: a live voice call where each participant hears the whole conversation before its turn. Three things change when models stop answering in isolation: Positions move. In independent answers, nothing ever changes its mind. In a shared conversation you see it happen: model B starts certain, hears model A's constraint it hadn't considered, and walks back its own suggestion. That retraction is the highest-signal moment in the whole session — it tells you which consideration actually mattered, something five parallel chats never show you because they never had to disagree. Errors get caught by the room, not by you. One model hallucinating an API flag tends to get corrected by another seat, roughly the way a colleague says "that wasn't in the docs last I checked." It is not a substitute for verification, but it beats you being the only reviewer of five confident paragraphs. The format forces brevity. Spoken turns are short. Counter-intuitively, that is a feature for decisions: you get "no, because X" instead of four thousand words of "it depends." For long-form analysis, written councils still win — a voice room is a debate club, not a research library. Running the room is a skill. What we have seen work across hundreds of sessions: Not a benchmark — confident is not correct, and a model agreeing with you is not evidence. Not a code reviewer; spoken turns are too short for long proofs. Not a legal or financial board. The niche where a live multi-model room genuinely beats both single-model chat and multi-tab comparison is decisions with trade-offs: architecture calls, pricing questions, "ship or polish" arguments, rehearsing a pitch against investors who interrupt. Whatever tool you reach for, the takeaway is architectural, not product-specific: the unit of AI leverage is moving from the answer to the argument between answers. Build your workflow around getting that argument — cross-review, ranked dissent, or just a room where the models can hear each other — and the model you pick matters a lot less than it feels like it should.