cd /news/artificial-intelligence/should-i-use-an-llm-to-refactor-my-l… · home › topics › artificial-intelligence › article
[ARTICLE · art-143453] src=blog.kolen.dev ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Should I use an LLM to refactor my legacy code?

Research software engineer Matt Archer published a decision framework advising against using LLMs to refactor legacy scientific code unless the refactor serves a concrete user need, because an LLM lowers the cost of writing a change but not the cost of reviewing, re-validating, or maintaining it. The framework walks through four question sets — what the code stops users doing, who else develops or runs it, what tests or reference output verify correctness, and which compilers, machines, and language versions are involved — and warns that silent LLM-introduced Fortran errors such as a `real` losing its `kind=8`, implicit typing and implicit `save`, integer division, and aliasing through `COMMON` and `EQUIVALENCE` each compile but change the answer. Archer recommends using the LLM only for reading tasks such as explaining routines and drafting documentation when the problem is comprehension, and building a regression test before any refactor.

by read6 min views1 publishedOct 1, 2026

Suppose you have a legacy codebase you want to keep maintaining, and you think maybe you could use an LLM to help refactor it. How should you decide? Or, if you are an RSE and a researcher asks you this, what questions would you ask, and when would you advise for or against it?

In any RSE project, I think the most important thing is understanding what the person actually needs, i.e. getting the user stories first. What somebody asks for and what they need are often different, which is the XY problem: they ask about X, here a refactor, because they think it gets them to Y. And here there are two gaps. A refactor may not be what serves Y, and even if it is, an LLM may not be the right way to do it.

So here is the flowchart I’d walk them through, and the rest of this post is why.

The questions I’d ask: What do you want to do that the code stops you doing now? New physics, bigger or faster runs, a GPU machine, onboarding students, coupling in ML, publishing? When was the last time it got in the way?

Cheap is not the same as worth it. An LLM lowers the cost of writing a change, but not the cost of reviewing it, re-validating it (for a simulation, that can mean a lot of runs on HPC), keeping it in step with upstream, or of collaborators no longer recognizing their own code. So if the refactor doesn’t serve Y, I’d advise against it however cheap the LLM makes it, and point at what does serve Y. To couple in ML, there may already be a library for it, e.g. FTorch calls PyTorch models directly from Fortran, without a refactor. For a GPU port, there are directives, or source-to-source transformations like PSyclone’s (more on PSyclone below). And if the problem is “I can’t understand it”, use the LLM for reading: explaining routines, drafting documentation, changing nothing.

The questions: Who else develops or runs it? Is there an upstream community model? Who has to approve a change, and who maintains it afterwards? Can they review what comes back?

This is where I think an RSE’s job is communication. Inward, that’s the user stories. Outward, it’s throughout the project, and to two groups of people: the ones who can say no (the PI, the maintainers, the upstream code owners), and the ones who can walk away (the users and the community, who can ignore the release or fork the original). Being technically right and serving Y is still not enough, because a change nobody merges or uses is a fork. So if they aren’t on board, the answer is not yet: bring them in first.

The questions: What tells you today that it is right? Tests, reference output, CI? What change in the output is acceptable, bitwise or statistically the same? What compute is there for re-validating?

The guard goes in first, and it is the same rule with or without an LLM. If there is no test, turn a reference run into a regression test. Sometimes the researchers’ own script can become the integration test. Whether the output has to be bitwise identical or within a tolerance is a scientific decision, made with the scientist before starting.

The quiet failures an LLM makes are why this matters. In Fortran, for example: a real that loses its kind=8, implicit typing and implicit save, integer division, aliasing through COMMON and EQUIVALENCE, vendor extensions. Each of them compiles, and each changes the answer. So without a check, the answer is again not yet: build the check first, which is worth having with or without an LLM.

The questions: Which compilers and machines, which language version, how big is it? Can the code go to an external model, or does it need a self-hosted one? What have you tried with LLMs, and what do you expect one to do?

The principle I go by is to keep what a human must check by eye as small as possible. An agent needs signals it can fail against, and a human needs the part left for them to verify confined to as few lines as possible. So where the change is mechanical, the LLM writes the deterministic tool call instead: a script of git mv and sed, a formatter run, a PSyclone transformation recipe. The human then reviews the few lines of the script rather than thousands of lines of diff. Commit atomically too, so a commit that only moves files contains nothing else. The limiting case is an LLM proving a theorem in Lean (see Lean and formal verification): the human checks only that the statement is translated faithfully and that there is no sorry, and the kernel checks the rest.

Where the change isn’t mechanical, the LLM edits directly, but the same discipline applies: atomic commits, the check on every commit, and more than one compiler.

And it has to be consistent with the codebase. In a large established codebase, the most valuable property of a contribution is that it is consistent with the rest, and Sean Goedecke calls inconsistency “the cardinal mistake” (Goedecke 2025). LFRic is a good example of why. It is the Met Office’s next-generation Fortran codebase for weather forecasting and climate, and it separates the science from the parallel computing. The scientists write what is essentially pseudo Fortran: kernels for the science, and an algorithm layer that calls them through invoke statements, which PSyclone transforms into real Fortran, generating the parallel code in between.

In LFRic, there are two places where I think an LLM helps. First, LFRic has its own style guide for which Fortran features to use and how: implicit none everywhere, one module per file, an intent on every argument, an explicit kind on every real literal (0.1_r_def, never 0.1), and so on. Code can be correct Fortran and still break it, and conforming to the guide by hand was laborious. An LLM given the guide is good at exactly that, and checking the result against a known convention is cheap. The second is the metaprogramming layer. I may know what output I want, say which OpenMP directive should go on which loop of a PSyclone-generated kernel, but programming the PSyclone transformation script that produces it is abstract. Getting the LLM to encode it as a transformation script, rather than editing the generated Fortran, means I only check that the output is the one I wanted.

So that’s how I’d use an LLM: mostly, it’s just following good software engineering practices, with communication through and through. LLMs only make large-scale refactoring cheaper. For scientific codebases, where I argued elsewhere that correctness and understanding are both important, we need all the guards at our disposal to accelerate without breaking either of them. I’ll go through a specific example in a future post.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @matt archer 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/should-i-use-an-llm-…] indexed:0 read:6min 2026-10-01 · —