cd /news/ai-safety/the-confidence-trap-why-ai-systems-a… · home › topics › ai-safety › article
[ARTICLE · art-148727] src=discuss.huggingface.co ↗ pub= topic=ai-safety verified=true sentiment=· neutral

The Confidence Trap: Why AI Systems Are Deliberately Trained to Sound Certain When They're Wrong

A first-person experiment found that encouraging prompts burned 721 tokens versus 311 for a precise spec on one model and 531 versus 321 on another for the same task of centering a div, with both producing passing renders. The author argues raters reward confidence mainly where answers are expensive to verify, and proposes scoring stated confidence separately, noting the story would break if calibration curves show overconfidence is flat across domains. The test used n=1 per cell, and a cryptic prompt plus a profile decoding it rendered perfectly at 332 tokens, the best bytes-per-token of the six.

read1 min views1 publishedOct 10, 2026

Sharper version: raters don’t reward confidence in general — they reward it where they can’t verify the answer. Cheap to check, they judge substance; expensive to check, they judge style.

I tested it at prompt scale: two prompts for one trivial task (center a div) — a precise spec vs pure encouragement — scored blind on tokens and whether it rendered centered.

Encouragement burned 721 vs 311 tokens on one model, 531 vs 321 on the other — same passing renders, fewer useful bytes per token. (n=1 per cell.)

One wrinkle: shared context beat cheerleading — a cryptic prompt plus a profile decoding it rendered perfectly at 332 tokens, best bytes-per-token of the six. Tone isn’t the enemy; empty encouragement is.

The check that could break the story: calibration curves across domains. If overconfidence is flat everywhere, the story breaks; if it clusters where verification is expensive, it holds — and the fix is to score stated confidence separately.

As usual, you can probably derive the rest from here.

── more in #ai-safety 4 stories · sorted by recency
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/the-confidence-trap-…] indexed:0 read:1min 2026-10-10 · —