cd /news/large-language-models/optimismbench-forecasting-bias-and-t… · home topics large-language-models article
[ARTICLE · art-79767] src=machinebrief.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment

A new study from arXiv introduces OptimismBench, a benchmark that detects directional bias in large language models' probability judgments by using inverted pairs to measure asymmetry between P(success) and P(failure) without ground truth. Testing 16 models from 8 providers, the study found that 14 are optimistic, with pessimism only in Anthropic's frontier tier, and that post-training sets the bias sign across 11 matched base-versus-chat pairs. The pattern persists across ablations, and model identity dominates language, with inter-model variance at 4.7x inter-language variance in a 17-model, 6-language comparison.

read1 min views1 publishedJul 30, 2026

arXiv:2607.26981v1 Announce Type: new Abstract: Large language models are increasingly used as decision aids whose probability judgments shape downstream choices. Whether those judgments carry a systematic directional tilt has been hard to detect: calibration metrics aggregate unsigned errors, and naturalistic uncertainty offers no ground-truth probability. When an LLM rates a startup's success at 70% but its failure at 15%, the missing 15 points expose a distortion no aggregate score flags. We introduce OptimismBench, which detects directional bias with inverted pairs: each scenario elicits both P(success) and P(failure), and asymmetry between the two framings yields a signed bias score without ground truth. Across 16 models from 8 providers, fourteen are optimistic; pessimism appears only in Anthropic's frontier tier. Eleven matched base-versus-chat pairs across four families show post-training sets the sign of the bias, with opposite shifts in different families. The pattern survives prompt, temperature, perspective, and self-debiasing ablations. A seventeen-model six-language comparison further shows model identity dominates language, with inter-model variance at 4.7x inter-language variance. We release 3,870 items across 10 languages for per-model directional-bias auditing. When alignment makes a model more helpful, it also tilts its probabilities; downstream pipelines inherit the tilt by default.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/optimismbench-foreca…] indexed:0 read:1min 2026-07-30 ·