{"slug": "grow-the-harness-not-the-context", "title": "Grow the Harness, Not the Context", "summary": "A September 22, 2026 arXiv paper by Laizhen Li introduces Growing Harness, a failure-guided training paradigm that learns an LLM agent's harness as reusable executable code from a strategy-free scaffold, cutting LLM calls by 76.0-91.8% and deployed-agent inference cost by 74.4-98.6% versus a Tool-Calling agent. Across BrowseComp-Plus and WebArena-Verified with three deployment models from 4B to 120B parameters, Growing Harness achieved the highest mean success in five of six benchmark-model settings and trailed the best mean by 0.7 percentage points in the sixth. On WebArena-Verified its success held at 44.7-45.3% across model scales, while Tool-Calling fell to 6.7% with the 4B model.", "body_md": "# Computer Science > Artificial Intelligence\n\n  [Submitted on 22 Sep 2026 (\n\n[v1](https://arxiv.org/abs/2609.26760v1)), last revised 24 Sep 2026 (this version, v2)]\n# Title:Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents\n\n[View PDF](https://arxiv.org/pdf/2609.26760)\n\n[HTML (experimental)](https://arxiv.org/html/2609.26760v2)\n\nAbstract:Large language model (LLM) agents often handle streams of related tasks, yet standard harnesses repeatedly ask the model to reconstruct the same control decisions inside each task's context. We study whether task feedback can instead turn recurring control into reusable executable code, while reserving LLM calls for task-specific semantic reasoning. We introduce Growing Harness, a failure-guided training paradigm that learns the agent harness itself from a strategy-free scaffold that exposes fixed model and tool interfaces but encodes no task-solving controller. Function-level execution traces localize each failure to a bounded code surface, an optimizer repairs a window of failures jointly, and a success-first held-out gate rolls back repair sequences that harm prior capability. Accepted edits accumulate in one shared harness, allowing its control structure to emerge from task feedback. Across BrowseComp-Plus and WebArena-Verified with three deployment models from 4B to 120B parameters, Growing Harness achieves the highest mean success in five of six benchmark-model settings and trails the best mean by 0.7 pp. in the sixth. Relative to a Tool-Calling agent, it reduces LLM calls by 76.0-91.8% and deployed-agent inference cost by 74.4-98.6%. On WebArena-Verified, its success remains 44.7-45.3% across model scales, whereas Tool-Calling falls to 6.7% with the 4B model. Ablations show that trace-local edits, joint repair, and gate-based rollback each improve final success. These results show that persistent program growth can move recurring control out of model context and into low-cost code, yielding reusable specialist agents that remain effective with smaller deployment models.\n    \n\n## Submission history\n\nFrom: Laizhen Li [\n[view email](https://arxiv.org/show-email/a718786d/2609.26760)]\n\n**Tue, 22 Sep 2026 17:40:45 UTC (353 KB)**\n\n[\\[v1\\]](https://arxiv.org/abs/2609.26760v1)\n**[v2]** Thu, 24 Sep 2026 08:15:47 UTC (353 KB)\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer \n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers \n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps \n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations \n\n*(*[What are Smart Citations?](https://www.scite.ai/))\n# Code, Data and Media Associated with this Article\n\nalphaXiv \n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers \n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub \n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub \n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face \n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast \n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))\n# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower \n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender \n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))\n# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/grow-the-harness-not-the-context", "canonical_source": "https://arxiv.org/abs/2609.26760", "published_at": "2026-09-26 16:20:01+00:00", "updated_at": "2026-09-26 16:31:21.692595+00:00", "lang": "en", "topics": ["ai-agents", "large-language-models", "artificial-intelligence", "ai-research", "ai-tools"], "entities": ["Laizhen Li", "Growing Harness", "BrowseComp-Plus", "WebArena-Verified", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/grow-the-harness-not-the-context", "markdown": "https://wpnews.pro/news/grow-the-harness-not-the-context.md", "text": "https://wpnews.pro/news/grow-the-harness-not-the-context.txt", "jsonld": "https://wpnews.pro/news/grow-the-harness-not-the-context.jsonld"}}