cd /news/ai-agents/meta-duke-and-uc-davis-researchers-u… · home › topics › ai-agents › article
[ARTICLE · art-144628] src=cryptobriefing.com ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

Meta, Duke and UC Davis researchers unveil self-improving branches for agent harness optimization

A preprint from Meta, Duke University and UC Davis researchers reports that a Mixture of Self-Improving Branches approach to agent harness optimization raised Olympiad-level math accuracy from 46.0% to 62.0% with Gemini 3 Flash, a 34.8% relative improvement, without retraining the model. Posted September 29, 2026 as arXiv:2609.37834v1, the method splits harness search into specialized branches that evolve on separate development data and routes each task to the best branch head, also posting an 11.6% gain on Terminal-Bench 2.0 and 3.8% on SWE-bench Lite. The work builds on Meta-Harness, released March 30, 2026, and claims to beat it roughly six months later.

by read3 min views2 publishedOct 3, 2026
Meta, Duke and UC Davis researchers unveil self-improving branches for agent harness optimization
Image: Cryptobriefing (auto-discovered)

Meta official logo (public domain, Wikimedia Commons) — CryptoBriefing brand treatment

A new preprint splits harness search into specialized branches and routes each task to the best one, beating the earlier Meta-Harness system on math, terminal and coding benchmarks

Researchers from Meta, Duke University and the University of California, Davis have a new idea for making AI agents better. Leave the model alone and improve the scaffolding around it, using several self-improving teams instead of one.

Their preprint, titled “Mixture of Self-Improving Branches for Agent Harness Optimization,” reports a 34.8% relative improvement on Olympiad-level math reasoning. That moved accuracy from 46.0% to 62.0% with the Gemini 3 Flash model, without anyone retraining the model itself.

What a harness is, and why it matters #

More formally, an agent harness is the code framework wrapped around a large language model. It covers the prompts the model receives, the tools it can call, the context it gets to see and how its actions are executed.

Harness optimization is the practice of searching for a better version of that wrapper automatically. The new paper, published on September 29, 2026 as arXiv:2609.37834v1, builds directly on an earlier system called Meta-Harness.

How the branching approach works #

Meta-Harness, released March 30, 2026, had previously outperformed traditional methods across various benchmarks. The new work argues that a single search path leaves performance on the table.

So the researchers split the search into multiple specialized branches. Each branch evolves on its own, with a distinct subset of development data and its own policy for proposing changes.

The branches also learn from their history. Each one keeps the cases where it beats its siblings and refines its strategy based on how earlier attempts performed.

That raises an obvious question: who picks the specialist? The answer is a router. At deployment, it selects the best branch head for each incoming input.

AI, tech, and the markets they move—in one daily briefing.

Daily. Free. Join 34,000+ readers across crypto, finance, and policy.

The whole process runs only on development-set data. The authors report that the complementary harnesses delivered their gains without any access to the test set.

The benchmark numbers #

The headline result came on Olympiad-level mathematical reasoning. Accuracy climbed from 46.0% to 62.0% using Gemini 3 Flash, which the authors frame as a 34.8% relative improvement.

The gains carried over to agentic tasks. On Terminal-Bench 2.0, which tests agents working in a command-line environment, the system posted an 11.6% gain.

SWE-bench Lite, a software engineering benchmark, showed a 3.8% increase.

Who is behind the work #

The author list includes Haoyu Dong, affiliated with Meta and Duke University, and Zihao Lin, affiliated with Meta and UC Davis. Lizhu Zhang and Zhuokai Zhao are listed as co-last authors.

The paper is a preprint posted to arXiv. That means it has been shared publicly but has not necessarily cleared formal peer review.

What this means #

A jump from 46.0% to 62.0% on hard math, achieved without touching model weights, suggests meaningful improvements can come from engineering the environment a model operates in.

There are open questions worth watching. Running and maintaining multiple evolving branches plus a router likely adds complexity, and the paper’s results come from specific benchmarks and a specific model.

Meta-Harness arrived in March 2026 and was already beating traditional methods. Roughly six months later, its successor claims to beat Meta-Harness itself.

Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our

Editorial Policy.

── more in #ai-agents 4 stories · sorted by recency
── more on @meta 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/meta-duke-and-uc-dav…] indexed:0 read:3min 2026-10-03 · —