{"slug": "red-queen-hypothesis-a-new-way-forward-for-self-improving-ai", "title": "Red queen hypothesis – a new way forward for self-improving AI", "summary": "Researchers from the University of Cambridge, NVIDIA, and Flower Labs have developed the Red Queen Gödel Machine, a method for recursive self-improving AI agents that co-evolves the agent and its evaluator to avoid performance ceilings. In tests, co-evolved scientific paper writers achieved 1.78×–1.86× higher acceptance rates, and co-evolved graders reached 9% higher ground-truth accuracy. The method also reduces computational costs, according to the pre-print paper led by PhD student Alex Iacob.", "body_md": "Submitted by Rachel Gardner on Tue, 21/07/2026 - 14:54\n\n## At a time when there's keen public interest in AI that can make itself better, researchers here have tackled one of the central challenges affecting its development.\n\nThe research team, which includes collaborators from NVIDIA and Flower Labs, have come up with a new method for recursive self-improving AI agents to continue improving themselves (by repeatedly testing and enhancing their own code) without hitting the evaluation ceiling that they frequently encounter.\n\nTheir method also suggests a way of cutting the costs of the computational resource needed for the development of such AI agents.\n\nWhile agents can already improve themselves by editing their own code, testing variants, and keeping changes that perform better, this process is usually limited by a fixed evaluator, benchmark, or test suite. Once the agent has learned everything that fixed signal can distinguish, improvement slows or stalls.\n\n\"A self-improving agent can only get as good as the test that scores it,\" explains team member ** Alex Iacob**, a PhD student in the\n\n**under the supervision of Prof**\n\n[Machine Learning Systems Lab](https://mlsys.cst.cam.ac.uk/)**. \"The test does not merely measure progress, it defines it, so the efficacy of the test becomes a ceiling the agent cannot climb past.\"**\n\n[Nic Lane](/people/ndl32)Now the researchers have addressed this issue by having both the self-improving agent and the evaluator evolve together. \"Instead of improving an agent against a fixed test, we let the **evaluation evolve alongside the agent**,\" Alex adds. \"As the agent gets better, the evaluation also gets **harder**, and the bar keeps rising.\"\n\n### The Red Queen Gödel Machine\n\nAlex is the first author on the pre-print paper the research team has just uploaded to arXiv. ** The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators** shares the technical details of their work, along with some impressive results from using this framework across a number of tasks.\n\n*Figure: Agents and evaluators improving together. The Red Queen Gödel Machine searches through many possible versions of an AI agent, while also improving the evaluator that judges those agents. During each phase, the evaluator is kept fixed so progress can be measured reliably. At checkpoints, a stronger evaluator can replace the old one if it performs better on trusted ground-truth examples. Scores from the old evaluator are then removed, so the next phase is guided by the new, more demanding standard. This creates a self-improving loop in which agents and evaluators improve together, while the system remains anchored to reliable checks.*\n\nIn both scientific paper writing and reviewing, and (Maths) Olympiad-level proof writing and grading, the Red Queen Gödel Machine improved performance over previous self-improving AI agents.\n\n\"Co-evolved scientific paper writers reach 1.78×–1.86× higher acceptance rates under a diverse agent-as-a-judge panel,\" report the researchers in the paper, \"while co-evolved graders reach 9% higher ground-truth accuracy.\"\n\n### The Red Queen hypothesis\n\nThe framework's curious title references the Red Queen – a fictional character in Lewis Carroll's children's novel, *Through the Looking-Glass* – who famously tells Alice, the novel's heroine, that \"it takes all the running you can do, to keep in the same place.\" The Red Queen hypothesis, a hypothesis in evolutionary biology put forward in 1973, was named after the character as it proposes that species must constantly adapt, evolve, and proliferate in order to survive while pitted against other species that are also continually evolving.\n\nThis hypothesis has now been applied to AI.\n\n\"Instead of improving an agent against a fixed test, we let the evaluation evolve alongside the agent. As the agent gets better, the evaluation also gets harder, and the bar keeps rising.\"\n\nAlex Iacob, PhD student\n\nThe work was carried out by a team of researchers here and with the support of collaborators including NVIDIA, Departmental spin-out [ Flower Labs](https://flower.ai/), MBZUAI and Inria.\n\n### A path towards more capable open agent systems at lower cost\n\nAnd another key finding that emerged was that using open source models to carry out some of the work, in addition to computationally expensive tools like ChatGPT, could be a way of significantly lowering the costs of developing self-improving AI systems.\n\nIn one experiment by the researchers, to co-evolve AI reviewers and writers of scientific papers, the researchers used the ** NVIDIA Nemotron 3 Ultra, **alongside ChatGPT-5.5. In the paper-reviewing task, this hybrid setup approached the performance achieved by ChatGPT on its own, while reducing search-token costs by around 13 times.\n\nCo-author Professor Nic Lane says: \"This is a narrow result, and we should be careful not to overstate it. But it also indicates where this method can go. If open models can carry the bulk of the search while stronger frontier models guide the higher-level improvement process, then we have a plausible path toward much more capable open agent systems at far lower cost.\"\n\nAnd Daniel Burkhardt, Developer Relations Manager at NVIDIA, says: \"This effort shows how open models can play an important role in advanced agentic systems. NVIDIA Nemotron models are designed to be efficient and capable for reasoning and agent workloads, and this work shows how they can be used as part of a broader self-improvement loop.\"\n\nThe research team is at pains to point out that the work is preliminary, and that longer search horizons will be needed to understand how far the approach can scale. But the results point to a new class of self-improving systems in which agents and evaluators recursively bootstrap one another, opening a path toward more capable, efficient and open agentic AI.\n\nThe paper, The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators, is available on arXiv. The underlying method will be open sourced to support wider adoption, while the team continues to investigate how this self-improving approach performs across a broader range of AI tasks, including experiments being scaled significantly beyond these initial results.\n\n: Alex Iacob, Andrej Jovanović, William F. Shen, Daniel Burkhardt, Meghdad Kurmanji , Nurbek Tastan, Lorenzo Sani, Niccolò Alberto Elia Venanzi, Ambroise Odonnat, Zeyu Cao, Bill Marino, Xinchi Qiu, Nicholas D. Lane**The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators**", "url": "https://wpnews.pro/news/red-queen-hypothesis-a-new-way-forward-for-self-improving-ai", "canonical_source": "https://www.cst.cam.ac.uk/news/red-queen-hypothesis-new-way-forward-self-improving-ai", "published_at": "2026-08-16 20:01:13+00:00", "updated_at": "2026-08-16 20:10:34.852603+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-research", "ai-agents", "machine-learning"], "entities": ["University of Cambridge", "NVIDIA", "Flower Labs", "Alex Iacob", "Red Queen Gödel Machine", "arXiv", "Machine Learning Systems Lab", "Nic Lane"], "alternates": {"html": "https://wpnews.pro/news/red-queen-hypothesis-a-new-way-forward-for-self-improving-ai", "markdown": "https://wpnews.pro/news/red-queen-hypothesis-a-new-way-forward-for-self-improving-ai.md", "text": "https://wpnews.pro/news/red-queen-hypothesis-a-new-way-forward-for-self-improving-ai.txt", "jsonld": "https://wpnews.pro/news/red-queen-hypothesis-a-new-way-forward-for-self-improving-ai.jsonld"}}