cd /news/ai-agents/agent-harness-evolution-shapes-codin… · home topics ai-agents article
[ARTICLE · art-112244] src=arxiv.org ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Agent Harness Evolution Shapes Coding Agent Quality

A controlled longitudinal study of 35 sequential releases of the Qwen Code CLI, holding the underlying LLM constant, found that agent harness evolution significantly impacts coding agent quality, with release velocities exceeding two releases per day and thousands of issues within months across five major open-source harnesses. The study, led by Oussama Ben Sghaier, evaluated each release against 50 stratified SWE-bench Verified tasks and traced quality fluctuations to specific development patterns and architectural components, challenging the common practice of attributing regressions to the model rather than the harness.

read2 min views3 publishedAug 26, 2026
Agent Harness Evolution Shapes Coding Agent Quality
Image: source
[Submitted on 4 Jul 2026 (

[v1](https://arxiv.org/abs/2607.03691v1)), last revised 20 Jul 2026 (this version, v2)]# Title:Don't Blame the Large Language Model: How Agent Harness Evolution Shapes Coding Agent Quality

[View PDF](/pdf/2607.03691)

[HTML (experimental)](https://arxiv.org/html/2607.03691v2)

Abstract:Coding agents, autonomous systems that use large language models (LLMs) to resolve software engineering tasks, rely on agent harness: a middleware layer in between a developer and a large language model that orchestrates system prompts, tool execution, context management, and iterative reasoning loops. While these agent harnesses evolve at extreme velocities, no study has examined how this evolution affects agent quality (i.e., effectiveness and efficiency) over time. Practitioners regularly report quality regressions after agent harness updates, yet consistently attribute them to the underlying model rather than the harness itself. In this paper, we address this gap by conducting the first controlled longitudinal study that isolates the agent harness contribution. Unlike prior work that fixes the agent harness and varies the model, we fix the model and vary only the agent harness, evaluating 35 sequential releases to measure their impact on agent effectiveness and efficiency. We first empirically study the development and release evolution of five major open-source agent harnesses (i.e., Codex, Qwen Code, Gemini, OpenCode, and OpenHands), revealing extreme release velocities exceeding two releases per day and thousands of issues within months. We then perform a controlled deep dive into 35 sequential releases of the Qwen Code CLI, evaluating each against 50 stratified SWE-bench Verified tasks while holding the underlying LLM constant. We trace the resulting quality fluctuations to specific development patterns and architectural components, and illustrate our findings with concrete qualitative evidence linking individual pull requests to measured quality shifts.

Submission history #

From: Oussama Ben Sghaier [[view email](/show-email/cd11dfa0/2607.03691)]

**Sat, 4 Jul 2026 03:55:25 UTC (1,827 KB)**

[[v1]](/abs/2607.03691v1)**[v2]** Mon, 20 Jul 2026 18:11:41 UTC (5,229 KB)

Current browse context:

cs.SE

References & Citations

...

Bibliographic Explorer

(What is the Explorer?) Connected Papers

(What is Connected Papers?) Litmaps

(What is Litmaps?) scite Smart Citations

(What are Smart Citations?)# Code, Data and Media Associated with this Article alphaXiv

(What is alphaXiv?) CatalyzeX Code Finder for Papers

(What is CatalyzeX?) DagsHub

(What is DagsHub?) Gotit.pub

(What is GotitPub?) Hugging Face

(What is Huggingface?) ScienceCast

(What is ScienceCast?)# Demos Influence Flower

(What are Influence Flowers?) CORE Recommender

(What is CORE?)# arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

── more in #ai-agents 4 stories · sorted by recency
── more on @qwen code cli 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/agent-harness-evolut…] indexed:0 read:2min 2026-08-26 ·