When a coding agent fails, the usual instinct is to swap the model. The evidence points somewhere else. On SWE-bench Pro, Claude Opus 4.5 scored 45.9% in Scale AI’s standardized test harness and 55.4% running inside Claude Code. That’s a 9.5-point difference from the setup alone, when a model upgrade typically moves scores by 2 to 4 points.
Nine blockers on a feature that was two screens and a button