cd /news/ai-agents/the-same-model-three-agents-why-the-… · home › topics › ai-agents › article
[ARTICLE · art-145184] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

The Same Model, Three Agents: Why the Harness — Not the Model — Decides What Your AI Can Do

A growing body of 2026 research shows that an AI agent's harness — its loop, context management, tool interface and recovery logic — drives benchmark performance far more than the underlying model weights. Harness-Bench (arXiv 2605.27922), a 106-task diagnostic benchmark from Peking University and Qiyuan Tech spanning 5,194 execution trajectories, found completion and failure behavior varied with the harness rather than the model, while a separate study raised fail-to-pass rates from 28% to 49% on a 169-task SWE-bench Verified cohort using only context-shortening and stall detection. The Self-Harness project further showed agents rewriting their own runtime rules, yielding 33–60% held-out relative gains across MiniMax M2.5, Qwen3.5-35B-A3B and GLM-5 with model, tools and benchmark fixed.

by read4 min views3 publishedOct 5, 2026

One-paste order for Medium's new-story editor: title → body → diagrams → notebook link → checklist.

Swap the model in your AI agent and you expect it to get smarter. Swap the harness — the loop, the context manager, the tool interface, the recovery logic — and the same weights score 28% on one benchmark and 49% on the next. In 2026, the industry finally stopped pretending the model is the agent. The evidence says the wrapper is the variable.

Think of the language model as a brilliant chef. The harness is the entire kitchen around them: the recipe binder (context), the pantry and tools (tool interface), the workspace counters (state), the fire alarm and extinguishers (guardrails), the timer that says "stop plating and serve" (stopping rules), and the health inspector writing down everything that happened (tracing/audit).

Put that chef in a well-organized commercial kitchen and you get dinner service. Put the same chef in a dark room with a butter knife and you get nothing — not because the chef changed, but because everything the chef needed to act was missing. A harness is that kitchen: the system layer that decides what the model sees, which tools it can touch, how work continues after a failure, and when to stop. The model is the easy part. The kitchen is the hard part.

Every production agent is a loop. The model proposes, the harness executes. Concretely, a harness has seven jobs:

Miss any one of these and the model can't help you. That's the mechanical reason 40% of agentic AI initiatives are projected to be discontinued by the end of 2027 (Gartner, via September 2026 industry coverage): not because the models were weak, but because the harness wasn't built.

Harness-Bench put numbers on it (May 2026, arXiv 2605.27922). Researchers at Peking University and Qiyuan Tech built a diagnostic benchmark: 106 sandboxed, manually-reviewed tasks drawn from real agent-use patterns, 5,194 execution trajectories, same tasks and budgets across model-harness pairings. Result: substantial variation in completion, efficiency, and failure behavior — driven by the harness, not the weights. Their conclusion is the quote of the year: agent capability should be reported at the model-harness configuration level, not attributed to the base model alone. A companion analysis cited a 23.8-point gap between the best and worst configurable harnesses on shared tasks.

The ablations are brutal. Hold the model fixed, change only the wrapper, and scores swing wildly: GPT-4 Turbo solved 18.0% of 300 SWE-bench Lite tasks with a purpose-built codebase interface but only 11.0% driving a plain shell. A minimal versus full adapter on the same GLM 5.1 backbone scored 19.1% versus 73.4% on Claw-SWE-Bench. And in "Same Model, Different Harness" (arXiv 2608.26218, August 2026), a purely mechanical change — shortening older tool results as context filled, plus stalling detection — raised mean fail-to-pass fraction from 28% to 49% and complete solutions from 43 to 72 on a 169-task SWE-bench Verified cohort. No weights changed. The wrapper moved the score. (The honest counter-evidence: with a strong long-context model and a fully observable environment, simple scaffolding reached 50.8% on SWE-bench Verified — scaffolding that compensates for weak reasoning shrinks as models improve. Controls that carry accountability don't.)

Harnesses started editing themselves. Self-Harness (arXiv, reported September 2026) lets an agent rewrite its own runtime rules. Starting from a minimal harness on the DeepAgent SDK, it ran Terminal-Bench-2.0 tasks, detected recurring failure patterns, and wrote targeted patches — e.g., one model kept exploring dataset configurations until timeout, so the system wrote a "loop breaker" forcing it to stop after 50 tool calls and draft deliverables early. Held-out relative improvements: 33–60% across MiniMax M2.5, Qwen3.5-35B-A3B, and GLM-5 — model, tools, and benchmark held fixed; only the harness varied. The key design detail: an acceptance rule promotes only edits that improve failures without regressing other tasks.

The harness became the product. On September 10, 2026, OpenAI launched the Agents API in public beta: the managed Codex harness — session orchestration, automatic context compaction, recovery, durable sessions that run for days, sandboxes (OpenAI-managed, self-hosted, or via partners including Cloudflare, Modal, and Vercel), subagent parallelization (Ciridae's CTO reported 4× latency cuts), MCP and custom tool connections. One API call spins up a production-ready agent; you pay tokens and sandbox minutes, no harness fee. The strategic read, laid out when OpenAI open-sourced the Codex harness under Apache 2.0 in August 2026: give away the specification, sell the operated version. Gartner forecasts 40% of enterprise applications will include task-specific AI agents by the end of 2026, up from under 5% in 2025 — and the vendors are now competing on who runs the best kitchen, not who trains the best chef.

References: Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows (arXiv 2605.27922, May 2026, Peking University / Qiyuan Tech) · Same Model, Different Harness: Different Coding-Agent Results (arXiv 2608.26218, Aug 2026) · Self-Harness (arXiv, reported via VentureBeat, Sep 2026) · OpenAI Agents API public beta (announced Sep 10, 2026; Codex harness open-sourced Apache 2.0, Aug 2026) · Towards AI Fryday #8: 10 Agent Harnesses That Change What the Same Model Can Do (Sep 2026) · Oracle Developers: Building an agent harness that survives production (2026) · Gartner enterprise-agent forecasts via Sep 2026 coverage.

Download: [diagram-1-anatomy.png](https://danielsamfdo.github.io/blog/assets/diagram-1-anatomy.png)

Download: [diagram-2-same-model.png](https://danielsamfdo.github.io/blog/assets/diagram-2-same-model.png)

Download: [diagram-3-self-harness.png](https://danielsamfdo.github.io/blog/assets/diagram-3-self-harness.png)

Download: [diagram-4-harness-as-product.png](https://danielsamfdo.github.io/blog/assets/diagram-4-harness-as-product.png)

**Companion notebook:** the runnable tutorial for this post — [download it here](https://danielsamfdo.github.io/blog/assets/agentic-harness.ipynb) (open in Colab/Jupyter).
── more in #ai-agents 4 stories · sorted by recency
── more on @peking university 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/the-same-model-three…] indexed:0 read:4min 2026-10-05 · —