09:16
2026-08-05
vibeleaderboard.ai
artificial-intelligence
Evals are the failure mode: today's research says the score isn't measuring the risk
Today's AI research converges on a finding that evaluation scores often fail to measure actual risk: cost-sensitive policy text in code-review prompts shifts reported failure probabilities by 13-17 poโฆ