Evolutionary Refactoring: Let Agents Invent Codebase-Specific Patterns A team at CodeScene used an evolutionary agentic refactoring method to uplift 300,000 lines of complex C code in the Street Fighter III codebase in about three weeks, a task the author estimates would otherwise take five to six experts 12 to 18 months. The approach has coding agents build their own codebase-specific refactoring playbook, whose patterns act as matching rules that mechanically map codebase problems to generic design patterns and refactorings at scale. The most powerful aspect of AI is that it broadens our toolbox. Done right, agentic coding gives us strategic options that just weren’t conceivable before. One such example is uplifting entire codebases to remediate technical debt. An approach that just wasn’t economically or even practically viable before. It’s an urgent problem since technical debt has been growing for decades, and coding agents are likely to accelerate its accumulation. Interestingly enough, those same agents might also have given us a way out. Recently, we found a promising solution to this age-old problem. It came as a by-product of an agentic refactoring study https://codescene.com/blog/case-study-refactoring-at-scale-with-agents where we uplifted 300k lines of complex C code in a non-trivial domain: the Street Fighter III codebase. The gist of the method is to take an evolutionary approach to refactoring: let agents create their own playbook with codebase-specific refactoring patterns. Those patterns act as matching rules, allowing agents to mechanically map problems in the codebase to generic design patterns and refactorings. As the playbook grows, so does their ability to apply those transformations at scale. To see how it works and why it matters, let’s first cover the traditional challenges with the classic codebase rewrite. The changing economics of the infamous rewrite Not long ago, uplifting a legacy codebase used to be both time-consuming and risky. Many are the technical debt backlogs that never got priority. Even more numerous are the failed project re-writes. Joel Spolsky https://www.joelonsoftware.com/2000/04/06/things-you-should-never-do-part-i/ went as far as calling a rewrite “the single worst strategic mistake” any company could do. After three decades in the software industry I can only sympathize with Joel. I’ve seen more failed rewrites than I ever wished for. However, the economics of the big rewrite might have changed. In his essay, Joel motivates the claim with the cost of maintaining the existing codebase for years before the replacement is ready. Recent evidence indicates that this time window could now be shorter. Much shorter. How long does it take to refactor a complete codebase? The main reason that rewrites are so tempting is because the existing codebase is deemed to be beyond rescue. Adding new features is expensive, bugs start to mount, etc. The system can no longer support the business. The other option is to repair. Uplift the quality of the existing code by means of large-scale refactoring. This option is typically less attractive to engineers; the siren song of a greenfield project is always going to be more appealing. Although a safer path than the total rewrite, mass refactoring comes with its own challenges. Large-scale refactoring is a scarce expert skill. Refactoring in general, or even familiarity with fundamental software design principles, is far from universally distributed knowledge. To pull off a refactoring at the scale of 300k lines of C code, I’d estimate that you need 5-6 people for 12-18 months. And all of them need to be experts in not only C programming, but software design, testing, and the domain itself. That’s a rare find. Contrast those estimates with what agents can do. In our study, agents pulled off that task in just three weeks. Or, more accurately, within a couple of days; a significant part of the up-front effort was to design the process that allows agents to refactor safely at scale. I’ve covered the details of that study, conducted by Daniel Webb and Markus Borg, on my company blog https://codescene.com/blog/case-study-refactoring-at-scale-with-agents . In this article I want to take a deeper dive into how agents pulled that off by building their own playbook. AI all the way down Ideally, you want recurring agentic tasks to involve learning. In our refactoring study, learning meant designing a process where a refactoring playbook is built up iteratively. That way, the refactoring tasks become progressively more effective. The refactoring playbook was written by the agents themselves. Our tooling informed the AI about existing code smells and their location, but left the solution in agentic hands. The agents then attempted a multitude of code transformations. Some failed and were discarded. The survivors got documented in the playbook. In writing the playbook, the agents had to respect two important constraints: - Preserve behavior : All recipes in the playbook must be behavior preserving. Only so can we guarantee that a refactoring is, well, a refactoring as opposed to just changing code. - Improve quality : Agents used the 10-point scale Code Health metric https://codescene.io/docs/guides/technical/code-health.html as our deterministic measure of code quality. Any refactoring recipe must lead to an increased Code Health. The agents were allowed to roam freely within these constraints to explore what an AI can do on its own when proper feedback loops are in place. After a few days of intense token munching, the task was completed. A 300k rewrite moved all code to a perfect Code Health score of 10.0. And: we could still play the game. However, the most exciting discovery wasn’t that this worked. Although that would have been completely unbelievable just a year ago . Rather it was the innovations found in the playbook. Agentic innovation as a side-effect Once the complete system had been uplifted, I started to look into the playbook. The first few recipes looked familiar. Those were all the classics: Extract Function, Introduce Guard Clauses, etc. Just what I would have expected. But beyond those classics I found novelty: a long list of codebase-specific refactorings. Our process caused agents to discover recurring transformation shapes that they formalized and named. Here are a few examples: - Action Parameter handles duplicated control structures that differ mainly in which function they invoke. - Shared Index Range captures repeated loops that differ only in their start and end ranges. - Uniform Step Table turns heterogeneous function calls into a uniform, table-driven dispatch. At first this all appeared bogus. Why on earth would you need those narrow recipes? Well, simply because they bridge an important gap: they make generic knowledge actionable by contextualization. From generic patterns to local recipes To explain what I mean by contextualization of generic knowledge we need to look at an example. Take the Action Parameter pattern. It’s a refactoring for deduplicating code. This is how the recipe looks: As you can see, those recipes contain specific constraints as well as guidance on when the refactoring applies. They are all surprisingly specific at a level that no really human would ever consider. So how do they work? Take a look at the following code snippet from the Street Fighter game: The preceding two functions perform exactly the same operation. The only difference is that one operates in the X coordinate space, the other in Y. Everything else is identical. It’s a duplication of knowledge with the same process expressed twice. This design problem matches the pre-requisites for the AI-invented Action Parameter pattern. Here’s what the refactored code looks like: The resulting code is both simple and elegant: encapsulate the commonalities in a shared function, and parameterize with the behavior that varies. The two original functions turn into one-liners that just inject the X vs Y offset calculation respectively. At this point, you might be tempted to throw an HC Andersen fairy tale in my direction: isn’t this just the design pattern STRATEGY https://adamtornhill.substack.com/p/hidden-design-decisions-refactoring done with C function pointers? Is the emperor naked? Yes and no. I mean, giving our agents a generic instruction to “apply STRATEGY to deduplicate code” is way too broad. It’s not going to turn out well. The beauty of codebase-specific refactoring patterns like Action Parameter is how specific and narrow they are. Thanks to that narrowness and constraints, it becomes possible to mechanically map specific problems in your codebase to generic design patterns and refactorings. I like to think about it as: Generic refactorings and patterns offer the foundation, codebase-specific refactorings adapt to the specific problem at hand. The code transformation itself might be generic, e.g. refactor to STRATEGY, Extract Method, etc. The innovative step is how codebase-specific patterns make that generic knowledge actionable by contextualizing it. The patterns become like matching rules. That means agents can recognize and reuse those recipes at scale, which creates a completely new context. It’s a first step towards making agentic refactoring a predictable and mechanical process at scale. How cool is that? Speed enables new strategies In my first post on this substack https://adamtornhill.substack.com/p/welcome-to-code-for-humans-and-machines , I wrote that the future of software development is not less engineering, but rather engineering at a higher level of abstraction. The idea of building a codebase-specific refactoring playbook is a concrete example of one such higher-level engineering approach: refactoring knowledge can become specialized and evolutionary. Machines can afford that. Like all surprising insights, this idea seems obvious in retrospect. But it would have been practically impossible just a year ago, and consequently almost unfathomable to even consider. The enabler is agentic speed. As a startup founder, I am obviously intrigued by the speed-up that AI offers. But the real benefit was never a 2-3x or even 10x task acceleration. The real gain is the strategic options and opportunities that speed enables. The idea of refactoring 300k of complex code within days would have been science fiction just a few years ago, and not even particularly good or believable fiction. As someone who’s spent my career improving code quality and trying to devise better methods for tackling technical debt, this hits close to my heart. It’s the first 400x task I’ve experienced, and a glimpse of a future where technical debt might be a solved problem.