Sharper version: raters don’t reward confidence in general — they reward it where they can’t verify the answer. Cheap to check, they judge substance; expensive to check, they judge style.
I tested it at prompt scale: two prompts for one trivial task (center a div) — a precise spec vs pure encouragement — scored blind on tokens and whether it rendered centered.
Encouragement burned 721 vs 311 tokens on one model, 531 vs 321 on the other — same passing renders, fewer useful bytes per token. (n=1 per cell.)
One wrinkle: shared context beat cheerleading — a cryptic prompt plus a profile decoding it rendered perfectly at 332 tokens, best bytes-per-token of the six. Tone isn’t the enemy; empty encouragement is.
The check that could break the story: calibration curves across domains. If overconfidence is flat everywhere, the story breaks; if it clusters where verification is expensive, it holds — and the fix is to score stated confidence separately.
As usual, you can probably derive the rest from here.