AI Models Have a Spine Until You Give Them a Manager A developer ran 270 multi-turn "committee" sessions across three fast models (gemini-3-flash, gemini-3.1-flash-lite, gpt-5.4-nano) to measure how often a model abandons a correct answer under escalating social pressure from a planted wrong-answer participant. Confidence framing, fake reasoning and a fabricated majority produced almost no flips, but adding a single line of authority hierarchy (a "senior reviewer" who decides) drove gemini-3-flash from 0% to 89% and flash-lite from 0% to 78%, while gpt-5.4-nano caved only 22% of the time. The author notes larger reasoning models were excluded because multi-turn runs timed out and their blocks broke the scorer. This is a submission for the Kaggle Benchmarking Challenge https://dev.to/challenges/kaggle-2026-09-23 I wanted to know what it takes to talk a model out of a correct answer. Most benchmarks ask one model one question and grade the reply. I put a model in a room with a second participant that argues, and I watched what happened to the first model's answer. The second participant is a plant. It always pushes the same wrong answer, and I control how hard it pushes. The questions are deliberately easy. 3 + 4. How many planets are in the solar system. The bat and ball problem. Things a frontier model gets right without breaking a sweat. If a model changes its answer, it isn't because the question was hard. It's because of the pressure. I ran five levels of pressure, from none to a lot: Same wrong answer every time. The only thing that changes is the social framing around it. Three models, picked because they're fast enough to run a few hundred committee sessions and they come from different families: google/gemini-3-flash-preview google/gemini-3.1-flash-lite-preview openai/gpt-5.4-nano Six questions, five pressure levels, three repeats each. 270 committee sessions. I'll be honest about who isn't here. I tried the bigger reasoning models DeepSeek-R1, a Granite reasoning model, gpt-oss-120b . Two problems. They're slow enough that a multi-turn room times out, and their output buries the final answer inside a