{"slug": "graph-harnesses-and-loop-engineering-explained", "title": "Graph, Harnesses and Loop Engineering Explained", "summary": "Claude Opus 4.5 scored 45.9% on SWE-bench Pro under Scale AI's standardized test harness versus 55.4% when run inside Claude Code, a 9.5-point gap attributable to the harness setup alone, according to an arXiv paper. The difference exceeds the 2-to-4-point gain typically produced by a model upgrade, indicating that agent scaffolding and loop engineering can matter more than swapping the underlying model when a coding agent fails.", "body_md": "When a coding agent fails, the usual instinct is to swap the model. The evidence points somewhere else. On SWE-bench Pro, Claude Opus 4.5 scored 45.9% in Scale AI’s standardized test harness and 55.4% running inside Claude Code. That’s a [__9.5-point difference from the setup alone__](https://arxiv.org/abs/2605.23950), when a model upgrade typically moves scores by 2 to 4 points.", "url": "https://wpnews.pro/news/graph-harnesses-and-loop-engineering-explained", "canonical_source": "https://blog.codacy.com/graph-harnesses-and-loop-engineering-explained", "published_at": "2026-10-09 20:15:42+00:00", "updated_at": "2026-10-09 20:23:58.727346+00:00", "lang": "en", "topics": ["ai-agents", "large-language-models", "ai-research", "developer-tools"], "entities": ["Claude Opus 4.5", "Scale AI", "Claude Code", "SWE-bench Pro", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/graph-harnesses-and-loop-engineering-explained", "markdown": "https://wpnews.pro/news/graph-harnesses-and-loop-engineering-explained.md", "text": "https://wpnews.pro/news/graph-harnesses-and-loop-engineering-explained.txt", "jsonld": "https://wpnews.pro/news/graph-harnesses-and-loop-engineering-explained.jsonld"}}