Creuto is now an OpenAI Select Partner Read More

Software Architecture & Technical

AI refactoring legacy code: what 300K lines in 3 weeks proves

Agents refactored 300,000 lines of C in three weeks for $4,000. What AI refactoring legacy code proves, what it does not, and the setup it depends on.

AI refactoring legacy code: what 300K lines in 3 weeks proves

Two engineers pointed coding agents at a 300,000-line C codebase for three weeks and spent roughly $4,000 on tokens. They produced 2,903 commits across 726 files, 252,055 lines modified, and a Code Health score that moved from 5.6 to a perfect 10.0. If you are costing AI refactoring legacy code against a conventional modernisation project, the figure that decides whether the comparison holds is not the line count. It is that this team had a frame-by-frame correctness oracle telling them the instant behaviour changed. Your legacy system almost certainly does not have one.

The measured facts, published by CodeScene on 24 September 2026:

  • Codebase: Street Fighter III: 3rd Strike, 300,000 lines of C, taken from an open-source decompilation.
  • Duration and cost: three weeks of part-time work, roughly $4,000 in tokens.
  • Output: 2,903 commits, 726 files, 252,055 lines modified, 22 refactoring recipes and 82 supporting notes.
  • Quality signal: the CodeHealth MCP Server, giving agents a deterministic score to optimise against.
  • Correctness check: a replay-trace harness comparing the rollback state hash frame by frame.
  • Models: Claude Opus for the bulk of the work; smaller models plateaued on the remaining smells.

The part that is genuinely new is the playbook, not the line count

Agents rewriting a lot of code quickly is not the interesting result. What the agents built alongside the code is. Rather than applying a fixed catalogue of transformations, they accumulated a playbook, naming recurring shapes they found and recording the preconditions for each. Familiar entries appear — Extract Function, Guard Clauses, Parameter Object — but so do three that are specific to this codebase: Shared Index Range for repeated loops differing only in start and end ranges, Action Parameter for duplicated control structures differing mainly in which function they invoke, and Uniform Step Table for turning heterogeneous calls into table-driven dispatch.

Failed attempts went into the playbook too, including transformations that made Code Health worse. That is the mechanism doing the work: a deterministic score, a cheap way to test a transformation against it, and a written record of what did and did not move the needle on this particular code. It is a search loop with a scoreboard, not a model being clever in one pass.

AI refactoring legacy code depends on an oracle most systems do not have

A decompiled fighting game replays deterministically. You can run the same inputs and compare a state hash every frame, so any behavioural drift surfaces immediately and automatically. That is an unusually strong oracle, and it is the reason the agents could take 2,903 swings without a human reading every diff.

Adam Tornhill, CodeScene's founder, says as much in the write-up: automated tests and equivalence checks are, in his words, absolutely essential as safeguards for preserving behaviour. Which places the load on exactly the thing unhealthy codebases tend to lack. The absence of a trustworthy behavioural check is usually why a legacy system has not been refactored, not an incidental gap in the tooling.

The team's own selection process makes the point. Daniel Webb said they had considered a Gov.UK marine licensing codebase he had worked on, but judged it too healthy to be useful for the research that follows, and chose the game partly because they play it. The codebase was picked for the study, and the study needed a bad starting score and a perfect oracle in the same repository. That is a rare combination in legacy application modernization work, where the systems that most need the uplift are the ones nobody dares change.

The criticisms worth taking seriously

Reaction split along what the result proves rather than whether it happened, and the sceptics are making engineering arguments, not rhetorical ones.

Konrad Otrębski asked, in the discussion under Tornhill's post, whether the work was merged, whether it arrived as one enormous merge or many, and whether this was an experiment on open-source code rather than production code earning money. Webb answered the first two directly: merged to main on a fork, through 54 pull requests. Otrębski's follow-up set a higher bar — offer the same refactoring to a project like Grafana, with merging to master as the definition of done.

Denis Baltor went after the recipes themselves, arguing that DRY concerns duplication of knowledge and intent rather than identical lines of code, and that recipes defined by loops differing only in ranges may be collapsing the two. Asko Nõmm raised the measurement problem: because Claude Code and Codex are tuned to their own models, it is unclear how much of the result measures the model and how much measures the harness, and cost would vary the same way. He also noted that architecture remains unmeasured, so code can score healthy while fundamental design problems surface later.

Webb himself left the sharpest question open. Asked whether the harness caught subtle frame-timing regressions, he said there might be no definitive answer, because it ran as a pre-commit hook and some failures were fixed without being observed. On diff sizes he asked, rather than asserted: if you are not familiar with the code and no longer review every line, how big can a diff be? That is the verification bottleneck stated by one of the people who lived it.

Two of the headline numbers are forecasts, not measurements

CodeScene projects a roughly 70% reduction in AI-induced defects and roughly 45% less token waste from the uplift. Both are extrapolated from its earlier research, not measured in this project. The $4,000 and the three weeks were measured. The 12-to-18-month pre-AI comparison is CodeScene's own estimate of what a similar project would have taken, not an observed alternative run.

The study that follows is where the interesting evidence will come from: two functionally equivalent versions of the same system, one at Code Health 5.6 and one at 10.0, with students at Lund University implementing features in both using frontier models and comparing cost and quality. That design can actually test the claim. This case study cannot.

Can agents refactor legacy code safely in your codebase?

Only with the three conditions this project happened to have, and they are worth naming as a checklist rather than a conclusion. A deterministic behavioural check that runs cheaply and often. A quality metric the agents can optimise that is not the same thing as the tests. And review capacity sized for a stream of small diffs rather than one large one.

In the systems we modernise, the first condition is the one that is missing, and building it is the project. That usually means characterization tests around the behaviour you cannot afford to change, recorded from production traffic rather than written from a specification nobody can find. It is slower and less photogenic than an agent committing 2,903 times, and skipping it is how a mechanical transformation becomes a silent behaviour change. Coverage alone will not tell you the check is sound either — mutation testing shows what generated tests miss.

The honest reading of this case study is narrow and still useful: where mechanical transformation is well defined, correctness is machine-checkable, and a score exists that correlates with something you care about, the cost of large-scale remediation has dropped by an order of magnitude. Where any one of those is absent, you have not bought a cheaper refactor. You have bought a faster way to generate diffs nobody can verify. Decide which of the three you are missing before you budget for the first one.

Frequently asked questions

Agents can refactor a legacy codebase safely only where a deterministic behavioural check exists. CodeScene's case study used a replay-trace harness comparing a state hash frame by frame. Without an equivalent oracle, agent refactoring produces diffs whose correctness nobody can verify at the speed they arrive.

The case study measured three weeks of part-time work, roughly $4,000 in tokens, 2,903 commits, 726 files touched, 252,055 lines modified, and a Code Health score moving from 5.6 to 10.0. The projected 70% defect reduction and 45% token saving were extrapolated from earlier research, not measured.

Record characterization tests from production traffic before you change anything, so the behaviour you cannot afford to break is pinned by observation rather than by a specification. Then check the tests themselves with mutation testing, because coverage alone does not prove a generated suite would catch a regression.

Daniel Webb, one of the two engineers, said the work was merged to main on a fork through 54 pull requests. It was not merged into an upstream production project earning revenue, which is the objection practitioners raised most often in the discussion under the announcement.

No. Asko Nomm made this objection directly: architecture remains unmeasured by the metric, so code can score healthy while fundamental design problems surface later. Code Health measures file and function level structure, which is a useful proxy for maintainability but not a verdict on system design.

The case study cannot separate them. Because Claude Code and Codex are tuned to their own models, the comparison mixes model capability with harness quality, and token cost varies the same way. CodeScene settled on Claude Opus after smaller models plateaued on remaining code smells.

Written by

Akash Mohapatra

Akash Mohapatra

Co Founder & Director

1 Oct 2026

·

7 min read

Share

LET'S CONNECT

Connect with Creuto!

Ready to take the first step towards unlocking opportunities, realizing goals, and embracing innovation? We're here and eager to connect.

We don't just aim to fit in – we strive to stand out. Experience the perfect blend of innovation, excellence, and trust that makes us truly unforgettable. Discover the difference with Creuto.

© 2026 Creuto All Rights Reserved