# Graph, Harnesses and Loop Engineering Explained

> Source: <https://blog.codacy.com/graph-harnesses-and-loop-engineering-explained>
> Published: 2026-10-09 20:15:42+00:00

When a coding agent fails, the usual instinct is to swap the model. The evidence points somewhere else. On SWE-bench Pro, Claude Opus 4.5 scored 45.9% in Scale AI’s standardized test harness and 55.4% running inside Claude Code. That’s a [__9.5-point difference from the setup alone__](https://arxiv.org/abs/2605.23950), when a model upgrade typically moves scores by 2 to 4 points.
