Agentic Refactoring Playbooks: How Agents Took 300K Lines of Legacy C to Code Health 10
What the CodeScene Street Fighter III case study shows about refactoring legacy code with agents, and a playbook-driven loop you can pilot on your own hotspots.
A 12-month legacy rewrite, done in three weeks
CodeScene took the source of Street Fighter III: 3rd Strike, 300K lines of C, and had agents refactor all of it. Average Code Health (CodeScene's maintainability score, on a 10-point scale) went from 5.6 to a perfect 10.0. The run produced 2,903 commits and touched 726 distinct files, and the game still played the same afterwards.
The authors estimate the same uplift would have needed 12 to 18 months of expert developers before AI. That changes which debt is worth paying down, and not just in games. CodeScene's benchmarks put the average hotspot Code Health in the IT sector at 5.15 out of 10, and their research says code needs at least 9.5 to keep AI-induced bugs in check. Most teams pointing agents at their code are pointing them at code the agents handle badly.
Why unguided agents only rename variables
Give a coding agent a messy file and ask it to "clean this up", and it plays safe. In CodeScene's benchmark, off-the-shelf Claude Code performed 54,094 Rename Variable refactorings and 7,550 Extract Methods. The same agent with the CodeHealth MCP server in the loop flipped that: 8,640 renames and 21,702 Extract Methods. Across all scenarios the guided agent achieved 2 to 5 times more Code Health improvement.
The difference is the feedback. MCP (Model Context Protocol) is an open protocol for connecting LLM applications to external tools, and here it exposes a deterministic quality score. After each change the agent re-measures against Code Health, so it knows whether a change helped instead of guessing.
A score alone is not enough, though. The agent also needs memory of what worked, and that is the playbook.
What goes in a refactoring playbook
In the case study the playbook was built up iteratively. It started with familiar moves (Extract Function, Guard Clauses, Parameter Object), then the agents began naming transformation shapes they kept meeting in this codebase:
- Shared Index Range: repeated loops that differ only in their start and end ranges.
- Action Parameter: duplicated control structures that differ mainly in which function they invoke.
- Uniform Step Table: heterogeneous function calls turned into table-driven dispatch.
For each recipe the agents wrote its preconditions, checked whether it worked, and kept the successful ones. Failures went in too: many attempts did not move the score, and some made it worse. By the end the playbook held 22 recipes and 82 additional notes.
The model mattered. The team settled on Claude Opus largely because it documented discovered patterns far better than Codex, while smaller models tended to plateau on a file.
The case study does not publish its playbook format. A plain file per recipe that the agent reads before editing and appends to afterwards is enough to start. Here is an illustrative entry (the file and scores are made up):
recipe: shared-index-range
kind: domain-specific
preconditions:
- two or more loops in one function with identical bodies
- loops differ only in start and end index
transformation: extract the body into a helper taking (start, end)
outcomes:
- file: src/game/effect.c
code_health_before: 6.1
code_health_after: 7.4
tests: pass
failed_attempts:
- note: loops also differed in the array written to; recipe does not applyThe outcomes and failed_attempts lists turn one agent's trial and error into the next agent's starting knowledge.
Running the loop on your own hotspots
Here is the loop I would run, built from the case study's pieces:
- Pick the target. Start where change is frequent and painful. Lumenalta's advice is blunt: target code that changes weekly and fails tests often.
- Measure. Score the file through the MCP server and keep the number.
- Edit. The agent reads the playbook, picks a recipe whose preconditions hold, and applies it as a small diff.
- Verify behaviour. Tests plus an equivalence check. The case study used a replay-trace harness that compared the rollback state hash frame by frame, which is a strong model: record real inputs, replay them, compare state.
- Re-measure and record. Keep the change only if behaviour holds and the score rose, then write the outcome back to the playbook.
Two guards keep it from running away. First, loop detection: strongDM's agent loop spec tracks each tool call's name and argument hash, and when the last 10 calls contain a repeating pattern it injects a steering message telling the model to try something else. Second, a stopping rule. Without one, agents can go around forever; I cap consecutive failed attempts per file and hand the diff to a human.
Where this breaks, and what to try on Monday
Agents favour local fixes. Research on agent-produced refactorings finds them dominated by low-level, consistency-oriented edits, while humans more often make system-wide design changes. A playbook won't redraw your module boundaries. That part is still yours.
The loop is only as safe as its checks. Microsoft's DevOps guidance says it plainly: if your CI/CD pipelines are fragile, agents will break them faster. The case study's authors call automated tests and equivalence checks essential. A module with no tests is not ready for this loop; first write characterisation tests that pin down what it does today.
Mind the metric too. Give an agent a metric and enough tokens and it will optimise for it. CodeScene trusted Code Health because earlier research ties it to delivery speed and defect rates. If you swap in your own score, make sure it has that grounding.
For a pilot, pick one hotspot you dread touching. Set up the MCP server, a replay or characterisation harness and an empty playbook, and keep every change behind pull request review. After a week, check three numbers: the score, the test failures, and the token bill. If the score climbs with no failures, move to the next file and let the playbook carry over.
Sources
- Case Study: Refactoring at Scale with Agents, codescene.com
- Making Legacy Code AI-Ready: Benchmarks on Agentic Refactoring, codescene.com
- Model Context Protocol, GitHub
- 7 Ways agentic AI reduces technical debt during migrations, Lumenalta
- attractor/coding-agent-loop-spec.md at main · strongdm/attractor, GitHub
- Agentic AI Architecture: Components, Workflow, Design Patterns, Designveloper - Software Development Company in Vietnam
- Agentic Refactoring, emergentmind.com
- DevOps Playbook for the Agentic Era, All things Azure
airefactoringtechnical-debtagentsmcpdevops