Google Research Open-Sources RRSI: AI Agents That Improve Their Own Harness Without Overfitting Google Cloud AI Research, with UNC-Chapel Hill, Stanford and Washington University in St. Louis, released RRSI (Regularized Recursive Self-Improvement), an Apache 2.0 research framework that lets an LLM agent rewrite its own harness — prompts, tools, memory, control flow and sub-agents — without changing model weights. RRSI regularizes the improvement loop to prevent overfitting, lifting Terminal-Bench 2.1 from 74.2% to 80.2% and SWE-bench Verified from 82.0% to 83.8%, with all 6 held-out splits improving and 2.42M policy tokens per trial versus 3.80M for unregularized evolution. The code requires Python 3.10+ and accepts any LiteLLM model string, defaulting to Claude Opus 4.8 on Vertex AI. Google Cloud AI Research , with UNC-Chapel Hill, Stanford and Washington University in St. Louis, has released RRSI Regularized Recursive Self-Improvement https://regularized-rsi.com/ . It lets an LLM agent rewrite its own harness: prompts, tools, memory, control flow and sub-agents. Model weights never change. RRSI constrains the improvement loop itself, so gains hold on benchmarks the agent never optimized against. Deployable? Yes, as a research framework. The code is Apache 2.0 https://github.com/google-research/rrsi , needs Python 3.10+, and accepts any LiteLLM https://github.com/BerriAI/litellm model string. Defaults assume Claude Opus 4.8 on Vertex AI. Why Self-Improving Harnesses Overfit Harness evolution loops propose edits, score them on a fixed evolve set and keep the winner. The same tasks are reused every round, so the loop can memorize them. The RRSI research https://arxiv.org/abs/2609.24972 names 3 failure modes: benchmark-specific fitting, noise chasing and complexity accumulation. Each one widens the gap between evolve-set scores and real transfer. How RRSI Works RRSI keeps every harness component editable. It regularizes how the search moves instead. Proposal side - Annealed edit budget: a cosine schedule lets early rounds bundle several edits. Late rounds allow a single attributable change. - Evidence-aware credit: each candidate is logged with its component, hypothesis, diff, score change and cost change. The proposer reads this ledger, so falsified ideas are not retried. - Structured exploration: when progress stalls inside the noise band, budget shifts to components the run never touched. Selection side - Leakage critic: rejects task names, entities, answers or benchmark-specific logic before any scoring. - Noise-adjusted floor: gains must clear the variance measured on the unchanged base harness. - Cost rule: extra inference tokens must be paid for by measured gain. - Pruning: components that stop producing gains become deletion targets. The research team frame these as analogies to classic regularizers. The edit budget maps to L0, pruning to Lasso L1 and the cost rule to Ridge L2 . Results Across 8 Benchmarks - Terminal-Bench 2.1 https://github.com/harbor-framework/terminal-bench evolve split : 74.2% to 80.2%. - SWE-bench Verified https://github.com/SWE-bench/SWE-bench never used for selection : 82.0% to 83.8%. - Out of distribution: JobBench https://github.com/Job-Bench/job-bench-eval +4.7, GDPval https://openai.com/index/gdpval/ +3.5 and APEX-Agents https://www.mercor.com/apex/apex-agents-leaderboard/ +3.7 points. - EngDesign https://github.com/AGI4Engineering/EngDesign evolve +4.9; Frontier-Eng https://github.com/Einsia/Frontier-Engineering +4.3 Medal points. - Harvey LAB https://github.com/harveyai/harvey-labs : +1.1 on the evolve split, +2.3 on its held-out split. All 6 held-out splits improved. With Gemini 3.5 Flash https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5/ as the policy, Terminal-Bench 2.1 rose from 64.6 to 78.7. SWE-bench Verified rose from 76.8 to 79.0. The harness is also lighter. On the agentic workspace instance, RRSI uses 2.42M policy tokens per trial. Unregularized evolution uses 3.80M. The abstract reports this as 30% fewer; the project page says 36%. RRSI vs Closest Competitors Scores come from Table 1 of the RRSI research paper. All methods share the same starting harness, policy, evolve split and candidate budget. | Feature | RRSI | Meta-Harness https://arxiv.org/abs/2603.28052 | AHE https://arxiv.org/abs/2604.25850 | TTHE https://arxiv.org/abs/2607.08124 | HarnessX https://arxiv.org/abs/2606.14249 | |---|---|---|---|---|---| | Core idea | Regularized proposal and selection | Agentic proposer over code, scores and traces of all prior candidates | Observability-driven loop; edits paired with verified predictions | Evolves harness during test time, no gold labels | Modular typed primitives, trace-driven adaptation | | Model weights | Frozen | Frozen | Frozen | Frozen | Frozen | | Cost rule and pruning | Yes | No | No | No | No | | Harvey LAB evolve score | 90.5 | 93.0 | 90.7 | 91.1 | 91.8 | | OOD average H0 = 39.7 | 43.6 | 40.6 | 39.2 | 38.0 | 39.7 | Per the RRSI research team. OOD average is the mean of JobBench, GDPval and APEX-Agents, computed from Table 1 https://arxiv.org/html/2609.24972v2 . Meta-Harness leads on the Harvey LAB evolve split. RRSI has the smallest evolve gain but the only OOD average more than 1 point above H0. Interactive Explainer How RRSI Regularizes Agent Self-Improvement Proposer Reads full edit ledger, shrinking edit budget Leakage critic Screens diff before scoring Evaluate Run on evolve set Gate Noise floor + cost rule Harness H Accepted, then pruning check