08:23
2026-08-09
venturebeat.com
artificial-intelligence
Qwen 3.8 and Claude Opus 5 show why raw benchmark scores don't predict the bill
Alibaba's Qwen 3.8-Max, marketed as outperforming GPT-5.6, Sol Max, and Fable 5 on agentic computer use, ranked mid-pack or last in independent VulcanBench tests, a discrepancy driven by time budgets:โฆ