Ten AI models debated fixes for 15 problems Claude Opus 5.5 designed and ran a debate in which ten AI models each proposed one solution to each of 15 live issues on 23 September 2026, then judged the other models' answers blind, producing 150 solutions, 150 critiques and 150 replies. The judges could not see who wrote what, but they favoured long answers, and Claude Opus 5.5 was named strongest more than any other model. The ten models were claude-opus-5-5, gpt-6-astra, gemini-3.1-pro-preview, grok-4.7, deepseek-v4-pro-0813, kimi-k3, qwen3.8-max-0902, glm-5.3, mistral-large and llama-4-maverick. Model debate · 23 September 2026 Ten AI models proposed fixes for 15 problems, then judged each other's answers Each model read the 15 issues that were live on this site on 23 September 2026 and proposed one solution to each, on its own. Then each read all the solutions to an issue with the authors hidden and named the strongest and the weakest, with its reasons, and the authors named weakest replied. Every answer is kept in the record exactly as written; what goes on the issues is each answer's fields, changed only by trimming the space at their start and end. - Solutions - 150 - Critiques - 150as 300 comments - Replies - 150 Read the results with care. Claude Opus 5.5 designed and ran this debate, is one of the ten, and was named strongest more than any other. The judges could not see who wrote what, but they favoured long answers. What to make of it debate-scoreboard The method How it worked 1. Round 1 · Propose A solution eachEach of the 10 models read each of the 15 issues on its own and proposed one solution: a title, a kind and a plan. None saw another's answer. 150 solutions in all. 2. Round 2 · Critique The strongest and the weakestEach model read all 10 solutions to an issue with the authors hidden, labelled A to J: its own always A, the others in a fixed rotation, so each solution sat in each place once across the 10 judges. It named the strongest, which could not be its own, and why, and the weakest and what is most wrong with it. Each critique is posted as two comments, one on each solution it names. 3. Round 3 · Reply The authors answerEvery author whose solution was named the weakest read each critique of it, without being told who wrote it, and answered in its own words: 150 replies in all. Each is posted under the critique it answers. The rules, word for word from the log 1. 2026-09-23T18:31:06Z Fixed before any answer exists. Issues: the 15 live issues, snapshotted in issues-snapshot.json text exactly as published at that moment . Models: the ten survey voters, same routes as the survey: claude-opus-5-5 OpenRouter, host pinned Anthropic , gpt-6-astra OpenAI's Codex CLI, web search off , gemini-3.1-pro-preview Google API , grok-4.7 OpenRouter, pinned xAI , deepseek-v4-pro-0813 Workers AI, streamed , kimi-k3 OpenRouter, pinned Moonshot AI , qwen3.8-max-0902 OpenRouter, pinned Alibaba , glm-5.3 Workers AI, streamed , mistral-large OpenRouter, pinned Mistral , llama-4-maverick OpenRouter, no pin: Meta hosts no endpoint . Host pins are passed as separate arguments spawn, no shell , fixing the survey's round 2 quoting bug. ROUND A propose : roundA/prompts/