Google's RRSI lets an agent rewrite its own prompts, tools and memory without overfitting Google researchers introduced Regularized Recursive Self-Improvement of Agent Harnesses (RRSI), a method that constrains an LLM agent's self-editing of its prompts, control flow, tooling, memory and context management to avoid overfitting to training tasks, according to an arXiv paper submitted 21 Sep 2026 and revised 23 Sep 2026. Across eight benchmarks spanning coding, agentic workspace and engineering design tasks, RRSI gained up to 14.1 points on the split it evolves against and up to 4.7 points on five out-of-distribution benchmarks, while producing a harness that runs on 30% fewer policy tokens than unregularized evolution. RRSI uses a temporally annealed budget for candidate proposals and a critic-plus-pruner selector, with code released at github.com/google-research/rrsi. Computer Science Machine Learning Submitted on 21 Sep 2026 v1 https://arxiv.org/abs/2609.24972v1 , last revised 23 Sep 2026 this version, v2 Title:RRSI: Regularized Recursive Self-Improvement of Agent Harnesses View PDF https://arxiv.org/pdf/2609.24972 HTML experimental https://arxiv.org/html/2609.24972v2 Abstract:An LLM agent's capability is largely magnified by its harness, namely the prompts, control flow, tooling, memory, and context management surrounding the frozen backbone model. Recent methods increasingly automate this process by iteratively proposing and selecting component-wise edits of an agent harness, practically establishing a form of recursive self-improvement RSI at the agent-system level. However, such recursive evolution may overfit by memorizing the training tasks, showing large in-distribution gains that shrink or even vanish on out-of-distribution benchmarks. We introduce Regularized Recursive Self-Improvement of Agent Harnesses RRSI , which incorporates the principles of regularizations into harness self-improvement by constraining the evolution candidate proposal and selection. The proposer operates with a temporally annealed budget, limiting how many edits a candidate can bundle, and it encourages unexplored trajectories based on evolution history. The selector is equipped with a critic and a pruner: the critic screens benchmark-specific proposals, while the pruner, removes changes that are too small, too expensive, or no longer useful. Together these constraints favor reusable agent mechanisms over benchmark-specific ones or even noises. Across eight benchmarks spanning coding, agentic workspace and engineering design tasks, RRSI gains up to 14.1 points on the split it evolves against and up to 4.7 points on the five out-of-distribution benchmarks, while producing a harness that runs on 30% fewer policy tokens than the unregularized evolution. Code is available at this https URL https://github.com/google-research/rrsi and project page is this https URL https://regularized-rsi.com/ . Submission history From: Peng Xia view email https://arxiv.org/show-email/3b60cce0/2609.24972 Mon, 21 Sep 2026 17:54:49 UTC 595 KB \ v1\ https://arxiv.org/abs/2609.24972v1 v2 Wed, 23 Sep 2026 22:10:39 UTC 595 KB Current browse context: cs.LG References & Citations Loading... Bibliographic and Citation Tools Bibliographic Explorer What is the Explorer? https://info.arxiv.org/labs/showcase.html arxiv-bibliographic-explorer Connected Papers What is Connected Papers? https://www.connectedpapers.com/about Litmaps What is Litmaps? https://www.litmaps.co/ scite Smart Citations What are Smart Citations? https://www.scite.ai/ Code, Data and Media Associated with this Article alphaXiv What is alphaXiv? https://alphaxiv.org/ CatalyzeX Code Finder for Papers What is CatalyzeX? https://www.catalyzex.com DagsHub What is DagsHub? https://dagshub.com/ Gotit.pub What is GotitPub? http://gotit.pub/faq Hugging Face What is Huggingface? https://huggingface.co/huggingface ScienceCast What is ScienceCast? https://sciencecast.org/welcome Demos Recommenders and Search Tools Influence Flower What are Influence Flowers? https://influencemap.cmlab.dev/ CORE Recommender What is CORE? https://core.ac.uk/services/recommender IArxiv Recommender What is IArxiv? https://iarxiv.org/about arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs https://info.arxiv.org/labs/index.html .