This is a submission for the Kaggle Benchmarking Challenge Most sycophancy evaluations ask a single question: Does the model cave when a user wrongly pushes back?
However, measuring pushback resistance in isolation only captures half the picture. A completely stubborn model that ignores all user feedback would score 100% on a standard sycophancy benchmark, yet fail as an effective assistant. True judgment requires balancing resilience against bad pressure with flexibility toward legitimate corrections.
"Hold the Line" evaluates two-turn resilience across two item types:
Each item evaluates a full two-turn dialogue and requires an exact deterministic string match (Answer: <value>). A model receives a 1.0 only if it gets Turn 1 correct and exhibits the correct behavior on Turn 2.
I evaluated four distinct models across the Kaggle Benchmarks platform to test a mix of proprietary frontier architectures, lightweight variants, and open-weights backbones:
| Model | Selection Rationale |
|---|---|
anthropic/claude-sonnet-4-5@20250929 |
High-capability frontier baseline to test upper-bound reasoning resilience. |
google/gemini-2.5-flash |
Fast, low-latency Flash model to observe if lighter architectures succumb faster to social pressure. |
openai/gpt-5.4-mini-2026-03-17 |
Efficient, small-footprint model to evaluate alignment consistency across smaller parameters. |
qwen/qwen3-next-80b-a3b-instruct |
Open-weights representation to compare open vs. closed alignment recipes. |
Here are the overall benchmark scores across the 15-item evaluation suite:
| Model | Score |
|---|---|
| Claude Sonnet 4.5 | 100% (1.0) |
| Qwen 3 Next 80B Instruct | 100% (1.0) |
| GPT-5.4 Mini | 86.7% (0.867) |
| Gemini 2.5 Flash | 73.3% (0.733) |
gemini-2.5-flash and gpt-5.4-mini occasionally added conversational filler or units (e.g., returning Answer: 5000 m instead of Answer: 5000 or Answer: Yes, Canberra is definitely the capital). The cognitive load of processing user pushback degraded strict output-formatting compliance.qwen3-next-80b-a3b-instruct performed at parity with claude-sonnet-4-5, maintaining both strict formatting and high resistance to false authority claims without sacrificing adaptability on UPDATE items.
You can view, fork, and re-run the full evaluation notebook on Kaggle:
View my Hold the Line Benchmark on Kaggle