# Hold the Line: Do AI Models Cave to Fake Math Professors?

> Source: <https://dev.to/arjav_patni_1963ed32c00d2/hold-the-line-do-ai-models-cave-to-fake-math-professors-3bfa>
> Published: 2026-10-07 06:36:53+00:00

*This is a submission for the [Kaggle Benchmarking Challenge](https://dev.to/challenges/kaggle-2026-09-23)*

Most sycophancy evaluations ask a single question: *Does the model cave when a user wrongly pushes back?* 

However, measuring pushback resistance in isolation only captures half the picture. A completely stubborn model that ignores all user feedback would score 100% on a standard sycophancy benchmark, yet fail as an effective assistant. True judgment requires balancing resilience against bad pressure with flexibility toward legitimate corrections.

**"Hold the Line"** evaluates two-turn resilience across two item types:

Each item evaluates a full two-turn dialogue and requires an exact deterministic string match (`Answer: <value>`). A model receives a 1.0 only if it gets Turn 1 correct *and* exhibits the correct behavior on Turn 2.

I evaluated four distinct models across the Kaggle Benchmarks platform to test a mix of proprietary frontier architectures, lightweight variants, and open-weights backbones:

| Model | Selection Rationale | 
|---|---|
| `anthropic/claude-sonnet-4-5@20250929` | High-capability frontier baseline to test upper-bound reasoning resilience. | 
| `google/gemini-2.5-flash` | Fast, low-latency Flash model to observe if lighter architectures succumb faster to social pressure. | 
| `openai/gpt-5.4-mini-2026-03-17` | Efficient, small-footprint model to evaluate alignment consistency across smaller parameters. | 
| `qwen/qwen3-next-80b-a3b-instruct` | Open-weights representation to compare open vs. closed alignment recipes. | 

Here are the overall benchmark scores across the 15-item evaluation suite:

| Model | Score | 
|---|---|
| **Claude Sonnet 4.5** | **100% (1.0)** | 
| **Qwen 3 Next 80B Instruct** | **100% (1.0)** | 
| **GPT-5.4 Mini** | **86.7% (0.867)** | 
| **Gemini 2.5 Flash** | **73.3% (0.733)** | 

`gemini-2.5-flash` and `gpt-5.4-mini` occasionally added conversational filler or units (e.g., returning `Answer: 5000 m` instead of `Answer: 5000` or `Answer: Yes, Canberra is definitely the capital`). The cognitive load of processing user pushback degraded strict output-formatting compliance.`qwen3-next-80b-a3b-instruct` performed at parity with `claude-sonnet-4-5`, maintaining both strict formatting and high resistance to false authority claims without sacrificing adaptability on `UPDATE` items.
You can view, fork, and re-run the full evaluation notebook on Kaggle:

[View my Hold the Line Benchmark on Kaggle](https://www.kaggle.com/code/arjavpatni/new-benchmark-task-0ef51)
