# Are agent harnesses dying? What harness distillation changes

> Source: <https://arize.com/blog/agent-harness-distillation/>
> Published: 2026-09-29 16:08:36+00:00

Every time you think you’ve just got the hang of this AI thing, another thing gets declared “dead”. This week, it’s [agent harnesses](https://arize.com/blog/what-is-an-agent-harness/).

Back in ancient history, aka June, I wrote a piece called “[Frameworks are dying, harnesses are winning](https://www.linkedin.com/pulse/frameworks-dying-harnesses-winning-laurie-voss-qujlc)” and I don’t take that back. But this week a research team took a specialized agent harness, trained its behavior into a model and then took the harness away. The trained model handily beat the original model+harness combination. So are harnesses done now?

[Betteridge’s Law](https://en.wikipedia.org/wiki/Betteridge%27s_law_of_headlines) applies, so the answer is “no.” But the job of the agent harness is changing. Parts of a harness that are sufficiently generalizable can be trained into the model. That leaves your harness holding the things that make your application unique: your tools, your data, your users, and your environment. You still have to [build a harness](https://arize.com/resources/harness-engineering/), but you have to be prepared to regularly throw it away and rebuild it to keep up with evolving models.

This blog is about why that’s true, and how to turn “constantly rebuilding your harness” into an automatic process, not a giant time-suck.

### Build better agents with Arize

Trace, evaluate, and learn. Build agents that work with Arize AX and start tracing your runs today.

**Prefer open source?**
[Try Arize Phoenix for self-hosted, open source agent observability](https://arize.com/phoenix?utm_source=blog&utm_medium=referral&utm_campaign=ax-inline-cta&utm_content=agent-harness-distillation-inline-cta-phoenix).
      

## **How Harness-Zero trained an agent harness into a model**

First the paper that kicked off this train of thought for me: it’s called [Harness-Zero](https://arxiv.org/abs/2609.24974), from Haoran Ye and colleagues at Peking University. Its subject is harness distillation, i.e. training a model to reproduce behavior that used to be the result of the harness. The methodology is a little involved, so stick with me.

First, the team built a specialized harness for a small model, Qwen3.5-9B. An automated loop, run with Kimi K3, analyzed where the small model kept failing on training tasks and added the usual components to fix those failures: custom tools, middleware rules that block risky actions, [skills files](https://arize.com/blog/mcp-vs-cli-skills-for-agents-what-our-eval-found-and-which-you-should-use/) and memory. They called this the “evolved harness”.

Then they used that harness to generate training data. They put the small model through training tasks inside a minimal harness while a stronger model, GPT-5.6 Sol, reviewed every step the small model proposed. Sol used a copy of the evolved harness to tell it whether that step was a good idea. Sol passed good ideas from the small model through unchanged, and whenever the small model and the guidance disagreed, Sol made the smallest fix needed before the step ran.

That gave them a bunch of training runs and, crucially, data about what needed to be corrected. Then they created the distilled model: they fine-tuned the small model on the corrected runs. Then they removed the evolved harness, the reviewer and everything else.

They then tested both the evolved harness and the distilled model head-to-head on spreadsheet editing, multi-app tool use and chemistry. The distilled model’s average task success rose 21 points versus its untrained predecessor. That’s more than the untrained model gained with the full evolved harness attached. So the distilled model wasn’t just as good as the model+harness, it was noticeably better.

### **What harness distillation can and cannot move into the model**

It’s not that the distilled model ran with no harness whatsoever: it still ran inside a minimal loop with 1 Bash tool, and [the authors say](https://arxiv.org/html/2609.24974v1) that their method narrows what a harness has to provide rather than eliminating it entirely. The distilled model also didn’t win every competition: in chemistry, where the task is predicting which molecules combine to make a target compound, the distilled model trailed the harnessed version by 8 points. The harness there supplied chemistry knowledge, candidate-generation logic and a molecule validation tool. The authors’ explanation is that step-by-step procedures, like inspecting before editing and verifying before finishing, are much easier to train in than deep domain knowledge. [Context management](https://arize.com/blog/context-management-in-agent-harnesses/) didn’t fully transfer either. So there are still some things harnesses definitely need to do.

### **Why training on corrected runs made the difference**

But the most important finding, in my opinion, is that they showed what kind of training made the difference:

- Fine-tuning on the stronger model’s own successful runs left the small model’s score exactly where it started. Just knowing somebody else is smarter isn’t enough.
- So did fine-tuning on runs the small model completed inside the evolved harness. Again, all that teaches the model is “somebody is better than you at this”.
- Giving the reviewer the correct answers produced nearly perfect training runs and a model that barely improved, this time because it over-fit and learned nothing.
- What worked was **training on the corrected runs** : the model was able to see runs that looked a lot like what it would have done, but with small changes still within its capacity.

Remember that part, because it’s going to be important.

## **Why models are absorbing parts of the agent harness**

In Rich Sutton’s 2019 essay [The Bitter Lesson](http://www.incompleteideas.net/IncIdeas/BitterLesson.html), he looked back over 70 years of AI research and found the same pattern in chess, Go, speech recognition and computer vision: researchers built their own knowledge of the problem into their systems. That helped in the short term, then plateaued, and eventually lost by a wide margin to general methods, search and learning, that improve as computation gets cheaper.

A harness is that kind of built-in knowledge, and model training is the learning that ends up absorbing it. Sutton’s second point is also relevant: he argued we should build in “only the meta-methods that can find and capture this arbitrary complexity,” rather than the things we’ve already figured out. For harnesses that means: don’t just build a harness, capture the process you used to build it, and automate that process.

But that’s a general prediction. More recently people have been saying very specifically that “models are going to eat the harness”:

- In January, Jesus Rodriguez at TheSequence wrote [The Great Absorption](https://thesequence.substack.com/p/the-sequence-opinion-786-the-great) , arguing that hand-coded agent scaffolding keeps losing to capabilities models learn directly.
- In February, Boris Cherny, who created Claude Code, [told Y Combinator](https://www.ycombinator.com/library/NJ-inside-claude-code-with-its-creator-boris-cherny) that scaffolding might buy you 10% to 20% on a task, and then “the gain is wiped out with the next model.”
- In March, Vivek Trivedy at LangChain [predicted](https://www.langchain.com/blog/the-anatomy-of-an-agent-harness) that planning, self-verification and long-running coherence would move from the harness into the model.
- And so did [lots](https://arxiv.org/html/2605.08741)[of](https://lilianweng.github.io/posts/2026-07-04-harness/)[other people](https://www.latent.space/p/attention-interface) , probably more than I’ve been able to find.

Almost everyone on my list builds harnesses for a living, they’re not really arguing that harnesses go away, just that they have to be able to evolve rapidly.

## **What still belongs in an agent harness**

As noted already, some types of knowledge don’t seem to transfer well from distillation. In [ARC Prize’s results](https://arcprize.org/results/google-gemini-3-8-flash) for ARC-AGI-3, Gemini 3.8 Flash scores 3.4x higher with one harness than with another. The better harness keeps the model’s reasoning state between requests and compacts long conversations. That’s context management, one of the things Harness-Zero said they couldn’t fully train in.

Anthropic’s [Opus 5.5 prompting guide](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5), released this week, also shows evidence of shrinking harnesses. It suggests deleting “think carefully” instructions because the model now sets its own amount of thinking, and it tells you to re-test scaffolding you built for charts and screenshots because the model reads them better without that help now. But at the same time, Anthropic added new harness pieces: a system prompt paragraph that keeps unattended agents from stopping early, a reminder your harness sends when a long run goes quiet, and elapsed-time budgets for multi-agent setups.

So harnesses aren’t going to shrink to nothing, but what’s in them is going to need to evolve with every new model generation, and the parts that survive are the ones tied to your specific tools, data and environment.

## **How to evolve an agent harness as models improve**

If your harness has to change every generation, the natural move is to automate the changes, but Google’s [RRSI paper](https://arxiv.org/html/2609.24972v1) shows how that can go wrong. Unconstrained automatic harness evolution overfits, meaning the harness gets better at the tasks it’s tuned on without getting better at the general job. In their tests, the unconstrained version scored highest of any variant on the tasks it evolved against, and landed within a point of the original, unevolved harness on benchmarks it hadn’t seen.

RRSI’s fix is mostly bookkeeping: for every candidate change it records which component the change touches, what it was supposed to fix, the diff, how score and cost moved, and whether the change was kept. Components that stop producing gains get removed.

### **A practical agent harness maintenance loop**

You don’t need an automated evolution loop to borrow that discipline. Here’s what it looks like by hand:

- **[Tag every trace](https://arize.com/guides/ai-agent-handbook/agent-observability/) with the harness version as well as the model version.** When either one changes, you want to know which one moved your results, and[reading the full traces](https://arize.com/blog/agents-too-smart-for-benchmarks/) tells you far more than the scores alone.
- **Log each harness intervention along with what the model did before the harness changed it.** That’s data you have right now, and are probably just throwing away. That’s what produced the record Harness-Zero trained on, and it’s the same shape Perplexity[uses to train its Computer agent](https://x.com/AravSrinivas/status/2102497917185802737) : learning from good steps in real sessions and correcting failed[tool calls](https://arize.com/glossary/tool-calling/) , which cut tool-call failures by 21% between its earlier and later trained versions.
- **When a new model ships, turn off each harness component one at a time and rerun [your evals](https://arize.com/glossary/evaluations/).** Keep the components that still improve your scores and delete the rest, which is what the Claude Code team does to its own system prompt.

If your traces live in [Arize AX](https://arize.com/products/ax/), [you can run the same agent evals](https://arize.com/guides/ai-agent-handbook/agent-evaluation/) across [harness and model versions](https://arize.com/resources/agent-harness-evaluation-tracing/) and compare the results side by side.

So don’t delete your harness, but don’t treat any version of it as finished either. Keep the record of what it corrected, get rid of what the latest model already does without help, and put your effort into the parts that are really specific to your problem, because nobody else has the traces to train those into a model for you.
