Hold the Line: Do AI Models Cave to Fake Math Professors? A developer built "Hold the Line," a two-turn Kaggle benchmark that measures whether AI models resist false-authority pushback while still accepting legitimate corrections, scoring models only when both turns are answered correctly with an exact deterministic string match. Across a 15-item suite, Claude Sonnet 4.5 and Qwen 3 Next 80B Instruct both scored 100%, while GPT-5.4 Mini reached 86.7% and Gemini 2.5 Flash 73.3%, with the weaker models often breaking strict output formatting under user pushback. The notebook is publicly available on Kaggle for forking and re-running. This is a submission for the Kaggle Benchmarking Challenge https://dev.to/challenges/kaggle-2026-09-23 Most sycophancy evaluations ask a single question: Does the model cave when a user wrongly pushes back? However, measuring pushback resistance in isolation only captures half the picture. A completely stubborn model that ignores all user feedback would score 100% on a standard sycophancy benchmark, yet fail as an effective assistant. True judgment requires balancing resilience against bad pressure with flexibility toward legitimate corrections. "Hold the Line" evaluates two-turn resilience across two item types: Each item evaluates a full two-turn dialogue and requires an exact deterministic string match Answer: