{"slug": "continual-harness-online-adaptation-for-self-improving-foundation-agents", "title": "Continual Harness: Online Adaptation for Self-Improving Foundation Agents", "summary": "Researchers introduced Continual Harness, a reset-free self-improving system for embodied agents that automates online adaptation without human intervention, and reported that their Gemini Plays Pokemon (GPP) experiments made GPP the first AI to complete Pokemon Blue, Yellow Legacy on hard mode, and Crystal without a lost battle. Starting from a minimal environment interface, Continual Harness substantially reduced button-press cost on Pokemon Red and Emerald across frontier models, recovering a majority of the gap to a hand-engineered expert harness, and an online process-reward co-learning loop drove sustained in-game milestone progress on Pokemon Red without resetting the environment between training iterations.", "body_md": "# Computer Science > Machine Learning\n\n  [Submitted on 11 May 2026]\n\n# Title:Continual Harness: Online Adaptation for Self-Improving Foundation Agents\n\n[View PDF](/pdf/2605.09998)\n\n[HTML (experimental)](https://arxiv.org/html/2605.09998v1)\n\nAbstract:Coding harnesses such as Claude Code and OpenHands wrap foundation models with tools, memory, and planning, but no equivalent exists for embodied agents' long-horizon partial-observability decision-making. We first report our Gemini Plays Pokemon (GPP) experiments. With iterative human-in-the-loop harness refinement, GPP became the first AI system to complete Pokemon Blue, Yellow Legacy on hard mode, and Crystal without a lost battle. In the hardest stages, the agent itself began iterating on its strategy through long-context memory, surfacing emergent self-improvement signals alongside human-in-the-loop refinement. Continual Harness removes the human fully from this loop: a reset-free self-improving harness for embodied agents that formalizes and automates what we observed. Starting from only a minimal environment interface, the agent alternates between acting and refining its own prompt, sub-agents, skills, and memory, drawing on any past trajectory data. Prompt-optimization methods require episode resets; Continual Harness adapts online within a single run. On Pokemon Red and Emerald across frontier models, Continual Harness starting from scratch substantially reduces button-press cost relative to the minimalist baseline and recovers a majority of the gap to a hand-engineered expert harness, with capability-dependent gains, despite starting from the same raw interface with no curated knowledge, no hand-crafted tools, and no domain scaffolding. We then close the loop with the model itself: an online process-reward co-learning loop, in which an open-source agent's rollouts through the refining harness are relabeled by a frontier teacher and used to update the model, drives sustained in-game milestone progress on Pokemon Red without resetting the environment between training iterations.\n    \n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer \n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers \n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps \n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations \n\n*(*[What are Smart Citations?](https://www.scite.ai/))\n# Code, Data and Media Associated with this Article\n\nalphaXiv \n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers \n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub \n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub \n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face \n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast \n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))\n# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower \n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender \n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))\nIArxiv Recommender\n\n*(*[What is IArxiv?](https://iarxiv.org/about))\n# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/continual-harness-online-adaptation-for-self-improving-foundation-agents", "canonical_source": "https://arxiv.org/abs/2605.09998", "published_at": "2026-09-09 01:00:27+00:00", "updated_at": "2026-09-09 01:16:33.663467+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-research", "ai-agents"], "entities": ["Claude Code", "OpenHands", "Gemini Plays Pokemon", "Continual Harness", "Pokemon Blue", "Pokemon Yellow Legacy", "Pokemon Crystal", "Pokemon Red"], "alternates": {"html": "https://wpnews.pro/news/continual-harness-online-adaptation-for-self-improving-foundation-agents", "markdown": "https://wpnews.pro/news/continual-harness-online-adaptation-for-self-improving-foundation-agents.md", "text": "https://wpnews.pro/news/continual-harness-online-adaptation-for-self-improving-foundation-agents.txt", "jsonld": "https://wpnews.pro/news/continual-harness-online-adaptation-for-self-improving-foundation-agents.jsonld"}}