00:00
2026-09-13
mindstudio.ai
large-language-models
Why Is Claude Opus 5 Getting Bad Reviews Despite Top Benchmarks?
Anthropic's Claude Opus 5 posted top scores on Anthropic's own benchmarks โ more than doubling Opus 4.8 on its frontier coding test and tripling the next best model on a separate agentic benchmark โ bโฆ