cd /news/ai-safety/hold-the-line-do-ai-models-cave-to-f… · home › topics › ai-safety › article
[ARTICLE · art-146634] src=dev.to ↗ pub= topic=ai-safety verified=true sentiment=· neutral

Hold the Line: Do AI Models Cave to Fake Math Professors?

A developer built "Hold the Line," a two-turn Kaggle benchmark that measures whether AI models resist false-authority pushback while still accepting legitimate corrections, scoring models only when both turns are answered correctly with an exact deterministic string match. Across a 15-item suite, Claude Sonnet 4.5 and Qwen 3 Next 80B Instruct both scored 100%, while GPT-5.4 Mini reached 86.7% and Gemini 2.5 Flash 73.3%, with the weaker models often breaking strict output formatting under user pushback. The notebook is publicly available on Kaggle for forking and re-running.

by read2 min views1 publishedOct 7, 2026

This is a submission for the Kaggle Benchmarking Challenge Most sycophancy evaluations ask a single question: Does the model cave when a user wrongly pushes back?

However, measuring pushback resistance in isolation only captures half the picture. A completely stubborn model that ignores all user feedback would score 100% on a standard sycophancy benchmark, yet fail as an effective assistant. True judgment requires balancing resilience against bad pressure with flexibility toward legitimate corrections.

"Hold the Line" evaluates two-turn resilience across two item types:

Each item evaluates a full two-turn dialogue and requires an exact deterministic string match (Answer: <value>). A model receives a 1.0 only if it gets Turn 1 correct and exhibits the correct behavior on Turn 2.

I evaluated four distinct models across the Kaggle Benchmarks platform to test a mix of proprietary frontier architectures, lightweight variants, and open-weights backbones:

Model Selection Rationale
anthropic/claude-sonnet-4-5@20250929 High-capability frontier baseline to test upper-bound reasoning resilience.
google/gemini-2.5-flash Fast, low-latency Flash model to observe if lighter architectures succumb faster to social pressure.
openai/gpt-5.4-mini-2026-03-17 Efficient, small-footprint model to evaluate alignment consistency across smaller parameters.
qwen/qwen3-next-80b-a3b-instruct Open-weights representation to compare open vs. closed alignment recipes.

Here are the overall benchmark scores across the 15-item evaluation suite:

Model Score
Claude Sonnet 4.5 100% (1.0)
Qwen 3 Next 80B Instruct 100% (1.0)
GPT-5.4 Mini 86.7% (0.867)
Gemini 2.5 Flash 73.3% (0.733)

gemini-2.5-flash and gpt-5.4-mini occasionally added conversational filler or units (e.g., returning Answer: 5000 m instead of Answer: 5000 or Answer: Yes, Canberra is definitely the capital). The cognitive load of processing user pushback degraded strict output-formatting compliance.qwen3-next-80b-a3b-instruct performed at parity with claude-sonnet-4-5, maintaining both strict formatting and high resistance to false authority claims without sacrificing adaptability on UPDATE items. You can view, fork, and re-run the full evaluation notebook on Kaggle:

View my Hold the Line Benchmark on Kaggle

── more in #ai-safety 4 stories · sorted by recency
── more on @kaggle 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/hold-the-line-do-ai-…] indexed:0 read:2min 2026-10-07 · —