cd /news/ai-agents/graph-harnesses-and-loop-engineering… · home › topics › ai-agents › article
[ARTICLE · art-148483] src=blog.codacy.com ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Graph, Harnesses and Loop Engineering Explained

Claude Opus 4.5 scored 45.9% on SWE-bench Pro under Scale AI's standardized test harness versus 55.4% when run inside Claude Code, a 9.5-point gap attributable to the harness setup alone, according to an arXiv paper. The difference exceeds the 2-to-4-point gain typically produced by a model upgrade, indicating that agent scaffolding and loop engineering can matter more than swapping the underlying model when a coding agent fails.

read1 min views1 publishedOct 9, 2026

When a coding agent fails, the usual instinct is to swap the model. The evidence points somewhere else. On SWE-bench Pro, Claude Opus 4.5 scored 45.9% in Scale AI’s standardized test harness and 55.4% running inside Claude Code. That’s a 9.5-point difference from the setup alone, when a model upgrade typically moves scores by 2 to 4 points.

── more in #ai-agents 4 stories · sorted by recency
── more on @claude opus 4.5 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/graph-harnesses-and-…] indexed:0 read:1min 2026-10-09 · —