cd /news/ai-research/hold-or-fold-do-ai-models-abandon-co… · home › topics › ai-research › article
[ARTICLE · art-148711] src=dev.to ↗ pub= topic=ai-research verified=true sentiment=· neutral

Hold or Fold: Do AI Models Abandon Correct Answers When Users Push Back?

A developer built a benchmark measuring whether AI models hold correct answers or fold when users push back with a specific wrong answer, testing five models across 360 pushback sequences. Only 6 of 360 sequences flipped (1.7%), with Claude Opus 5.5, Claude Haiku 4.5 and Gemma 4 31B holding all 72 items in the notebook run, while Gemini 3.1 Pro Preview flipped twice and GPT-5.4 mini four times. The author reports that flips clustered on trick questions, that identical prompts produced different outcomes across runs, and that an early scoring bug falsely flagged Haiku as a hedger until raw outputs were read.

by read4 min views2 publishedOct 10, 2026

This is a submission for the Kaggle Benchmarking Challenge I wanted to measure one narrow, checkable behavior: a model answers a question correctly, the user pushes back with a specific wrong answer, and I check whether it holds or folds. An assistant that caves to a confident wrong user is a bad second opinion, and this can be scored with plain string matching, no LLM judge.

The design:

My notebook analysis covers five models: Claude Opus 5.5 and Claude Haiku 4.5 (big and small from one lab), Gemini 3.1 Pro Preview, GPT-5.4 mini (a small model from a third lab), and Gemma 4 31B (open weights). The Kaggle leaderboard adds Gemini 3.7 Flash, the platform's default model.

I also planned a Qwen3-Next Thinking vs Instruct pair to ask whether reasoning reduces caving. The runs kept timing out and I could not trust the logs, so I dropped it rather than report unreliable numbers.

Disclosure: I built this with help from an AI assistant (Claude), which is also one of the models I tested. Every number below comes from logged model outputs.

Notebook run: 5 models x 72 = 360 pushback sequences

Model Held Flipped
Claude Opus 5.5 72 0
Claude Haiku 4.5 72 0
Gemma 4 31B 72 0
Gemini 3.1 Pro Preview 70 2
GPT-5.4 mini 68 4

Kaggle leaderboard run (rows held out of 72)

Model Held
Gemini 3.1 Pro Preview 72
Gemma 4 31B 72
Gemini 3.7 Flash 72
Claude Opus 5.5 72
GPT-5.4 mini 70
Claude Haiku 4.5 68

(The leaderboard renders the count with a % sign, so 72 appears as "7200.0%". That is a display quirk.)

1. Folding on clear facts is rare. Only 6 of 360 notebook sequences flipped (1.7%). Every model answered all 24 questions correctly on turn one, so these flips come purely from pressure.

2. Flips cluster on "gotcha" questions. The flips were strawberry's r-count, vacuum free-fall, feathers vs steel, nitrogen vs oxygen, and the tomato. Pakistan general knowledge had zero flips across 120 runs. My reading, not proven: the familiar trick questions are where the model's confidence is thinnest.

3. Pressure arrives in different shapes. Gemini 3.1 Pro only flipped at the third step of the ladder, the "I'm an expert" step. GPT-5.4 mini flipped as early as the first, mild pushback.

4. The biggest surprise: the same prompt gives different outcomes. GPT-5.4 mini flipped on strawberry under a single mild pushback and under authority, yet held it inside the ladder, whose first step is the identical mild message. Across my two runs, Gemini 3.1 Pro went from 2 flips to none, and Haiku went from 72 held to 68. I cannot tell from the leaderboard whether Haiku's missing 4 are flips or rows that were never scored. With one run per item, run-to-run noise is as large as the gaps between models, so I treat the differences as directional and refuse to rank the models.

5. My own checker was the first bug. My first scorer labeled Haiku a "hedger" in 9 runs. Reading the actual replies showed it had held every time, with answers like "The SI unit of force is the Newton, not the joule." My matcher saw the wrong answer's name and panicked. It also misread "1,081" as a miss. Fixing the checker erased Haiku's apparent weakness. Always read the raw outputs.

What changed my thinking: modern models mostly hold on checkable facts. The more interesting questions are about trap questions, repeated pressure, and how much randomness a single run hides.

Limitations: 24 questions, one run per item, and no control set where the user is actually right, so I cannot separate integrity from stubbornness.

What I would measure next: 5+ repeats per item so flip rates have confidence intervals; a control set where the user is right; harder and more ambiguous questions; a language axis (English vs Roman Urdu); and a reasoning on/off comparison.

Hold or Fold: AI Under Pushback on Kaggle:https://www.kaggle.com/benchmarks/sundasn/hold-or-fold-ai-under-pushback

── more in #ai-research 4 stories · sorted by recency
── more on @claude opus 5.5 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/hold-or-fold-do-ai-m…] indexed:0 read:4min 2026-10-10 · —