{"slug": "evoharnessrl-learning-self-evolving-runtime-harness-for-long-horizon-llm-agents", "title": "EvoHarnessRL: Learning Self-Evolving Runtime Harness for Long-Horizon LLM Agents", "summary": "Researchers introduced EvoHarness-RL, a framework that trains long-horizon LLM agents to learn runtime harness policies for constructing and coordinating external state, achieving 96.9% success on ALFWorld with a Qwen3-8B LLM. The method uses supervised fine-tuning and cost-aware GRPO to teach agents to selectively read, update, and consolidate Belief, Progress, and Experience (BPE) state, revealing dynamics of harness annealing and evolution.", "body_md": "# Computer Science > Machine Learning\n\n[Submitted on 5 Aug 2026]\n\n# Title:EvoHarness-RL: Learning Self-Evolving Runtime Harness for Long-Horizon LLM Agents\n\n[View PDF](/pdf/2608.05446)\n\n[HTML (experimental)](https://arxiv.org/html/2608.05446v1)\n\nAbstract:Long-horizon LLM agents increasingly rely on external execution support to maintain state, track progress, invoke tools, verify outcomes, and reuse experience across interactions. However, effective harness use raises two coupled challenges: state formation from noisy interaction traces and runtime control over external-state access. Existing agents usually handle both through prompts, heuristics, or domain-specific conventions, leaving the external workspace and its usage policy manually engineered. To address this, we study the problem of harness policy learning, where agents learn harness policies offline and deploy them to construct and update external harness state online during runtime task execution. We introduce EvoHarness-RL, which exposes Belief, Progress, and Experience (BPE) as policy-facing harness state. Supervised harness fine-tuning teaches the base agent the harness action space and how to construct useful external state, while cost-aware GRPO explores coordination policies to selectively read, update, and consolidate that state during long-horizon interaction. Instantiated on ALFWorld with a Qwen3-8B LLM, EvoHarness-RL reaches 96.9% success and reveals two key dynamics: harness annealing, where training internalizes recurring harness-use patterns into the model policy and shifts the agent from frequent harness calls toward selective external-state access, and harness evolution, where progress updates and experience consolidation refine the harness into a compact, task-adaptive state substrate. These results suggest that long-horizon agents benefit from trainable policies for constructing and coordinating with external harness workspaces, beyond simply adding stronger tools or larger memories.\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer\n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers\n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps\n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations\n\n*(*[What are Smart Citations?](https://www.scite.ai/))# Code, Data and Media Associated with this Article\n\nalphaXiv\n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers\n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub\n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub\n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face\n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast\n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower\n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender\n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))\nIArxiv Recommender\n\n*(*[What is IArxiv?](https://iarxiv.org/about))# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [ Learn more about arXivLabs](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/evoharnessrl-learning-self-evolving-runtime-harness-for-long-horizon-llm-agents", "canonical_source": "https://arxiv.org/abs/2608.05446", "published_at": "2026-08-10 21:28:07+00:00", "updated_at": "2026-08-10 21:41:39.479205+00:00", "lang": "en", "topics": ["machine-learning", "large-language-models", "ai-agents", "ai-research"], "entities": ["EvoHarness-RL", "ALFWorld", "Qwen3-8B", "GRPO"], "alternates": {"html": "https://wpnews.pro/news/evoharnessrl-learning-self-evolving-runtime-harness-for-long-horizon-llm-agents", "markdown": "https://wpnews.pro/news/evoharnessrl-learning-self-evolving-runtime-harness-for-long-horizon-llm-agents.md", "text": "https://wpnews.pro/news/evoharnessrl-learning-self-evolving-runtime-harness-for-long-horizon-llm-agents.txt", "jsonld": "https://wpnews.pro/news/evoharnessrl-learning-self-evolving-runtime-harness-for-long-horizon-llm-agents.jsonld"}}