A leading product engineering company, creating adaptive software solutions to improve operations, providing businesses with expert development services from across domain.
A leading product engineering company, creating adaptive software solutions to improve operations, providing businesses with expert development services from across domain.
You cannot refactor safely without knowing what the code does. Characterization tests legacy code by recording actual behaviour and hashing it.

The hardest part of modernising a legacy system is not writing the new code. It is proving the new code does what the old code did, when nobody can tell you what the old code did. Characterization tests legacy code problems this way round: instead of describing intended behaviour, you record actual behaviour — bugs included — and treat that recording as the contract.
The technique is old and well documented. What has changed is that generating the recording is now cheap, and that coding agents have made the question urgent rather than academic.
A characterization test makes no claim that the output is correct. It asserts only that the output is what it was before you touched anything. That distinction is what makes it usable on code nobody understands.
A workflow published this week makes it concrete, and the discipline is in the order of operations.
The hash is what makes this work in practice rather than in theory. A test suite tells you which assertion failed; a hash tells you instantly whether anything at all moved, across every case simultaneously, without anyone having decided in advance which outputs mattered.
It is also why the fixture matrix matters more than the number of cases. Coverage here is not about lines executed, it is about the input space: the combinations where two code paths could both apply, the values that only appear in production, the shapes the original author never anticipated but that customers send anyway. A matrix of forty carefully chosen inputs is worth far more than four hundred generated at random, because the interesting behaviour of legacy code lives in the corners where its rules overlap.
A question that comes up immediately: what if the recorded behaviour is wrong? It almost certainly is, in places. The answer is that fixing it is a separate, deliberate commit — you change the snapshot and the code together, and the diff on the snapshot file shows a reviewer exactly which outputs you intended to change. That is a far better review artefact than a behavioural change buried inside a refactor, where nobody can tell the intended change from the accidental one until a customer reports it.
Snapshot approaches fail for boring reasons, and the write-up is unusually good on them.
Serialise deterministically. Sort keys, fix the trailing newline. An unordered dictionary produces a different hash on a different run for no reason at all, and one false alarm is enough for a team to stop trusting the tripwire.
Watch parallel test workers. If several workers write the same fixture file, they clobber each other and the failure looks like behavioural drift. This one costs an afternoon to diagnose the first time.
Expect language-level changes. Unicode handling shifts across Python versions, and so do floating point formatting and dictionary iteration in other runtimes. A hash change on a runtime upgrade is real information — it means behaviour changed — but it is not your refactor.
The snapshot lives in the commit, which is the point. It is not a test someone runs before a release, it is a tripwire that fires on the change that broke it, next to the change that broke it.
One practical note on scope. This works on a module, not on a system. The observable inputs of an entire application are unbounded, and trying to characterise everything produces a matrix nobody maintains. Pick the module you are about to change, characterise its boundary, and accept that you are buying confidence about one seam rather than about the whole estate.
This is the part worth arguing about, because the obvious use is the wrong one.
The tempting move is to point a coding agent at a messy module and ask it to refactor. It will produce something plausible, reviewable, and very hard to verify — because the reviewer does not know the original behaviour either. You have generated a large diff against an unknown baseline and made the review harder rather than easier.
The productive use is at step two. Ask the model to suggest fixture cases you have not thought of: the collision between two rules, the empty file, the encoding that breaks the parser. Finding overlooked edge cases is exactly the kind of exhaustive, unglamorous enumeration a model is good at, and it is the step humans skip because it is boring.
Then the harness runs locally and the snapshot is the authority. The model proposes cases; it does not adjudicate whether behaviour changed. That division matters, and it is the same principle as scoring the trajectory rather than the final answer: the check has to be mechanical, because a model asked to confirm its own work will confirm it.
There is a real constraint to respect too. Do not build fixtures from live customer data, and be careful pointing this at security-sensitive parsers, where recording current behaviour may be recording the vulnerability. Characterization tests preserve bugs by design, which is their value during a refactor and a liability if you forget.
The economics are worth stating plainly, because this always reads as extra work to a sponsor watching a budget. Characterising a module takes perhaps a day. Not characterising it means every subsequent change is reviewed by someone reasoning about behaviour from the code alone, which is slower per change and unreliable in a way that surfaces as production incidents weeks later. The day is repaid on roughly the third change, and everything after that is profit.
It also changes what a partial migration means. With a tripwire in place you can stop after four extractions, leave the module half-modernised, and know the system still behaves identically — so the work can be paused for a quarter without risk. Without one, a half-finished refactor is a liability nobody wants to inherit, which is precisely why so many modernisation efforts are all-or-nothing and why so many of them are abandoned in the middle.
We get asked this most weeks now, usually about a system a decade old with no tests and nobody left who wrote it.
The honest answer is that the bottleneck was never writing the replacement. It was that any change to that system is unfalsifiable — there is no way to demonstrate the new version behaves like the old one, which is why these projects stall for years and why rewrites overrun. Generation being cheap does not touch that problem. It makes it worse, because now you can produce untrustworthy change faster.
Characterization tests attack the actual constraint. They convert "nobody knows what this does" into a committed artefact that fails loudly when behaviour moves, and once that exists an agent becomes genuinely useful — because every suggestion it makes is now checkable in seconds.
So the sequencing we would recommend on any modernisation engagement is: characterise first, refactor second, and let the model help with the characterising rather than the refactoring. That inverts what most teams try, and it is the difference between a migration that can be paused safely at any commit and one that has to complete or be abandoned. It is also why we treat this as the opening phase of legacy modernisation work rather than a testing afterthought, and why the same logic shows up in modernising systems without breaking the business.
If your system has no tests and a refactor is being discussed, the first week is not architecture. It is building the tripwire, and it is worth more than the first month of the rewrite — which is the sort of judgement any quality engineering engagement should be making before code is touched.
A test that records what code currently does rather than asserting what it should do. It makes no claim that the output is correct, only that it matches the behaviour captured before any change, which is what makes it usable on code nobody fully understands.
Inventory the observable inputs, build a fixture matrix covering edge cases and collisions, record the actual outputs to JSON and hash the file. Then extract one small function per commit and re-run the matrix, reverting if the hash drifts.
A hash tells you instantly whether anything changed across every case at once, without anyone deciding in advance which outputs mattered. Assertions only catch what someone thought to assert, which on unfamiliar code is the wrong set.
Non-deterministic serialisation is the main cause, so sort keys and fix trailing newlines. Parallel test workers writing the same fixture file clobber each other, and language-level changes such as Unicode handling across Python versions shift output legitimately.
Not safely on its own, because the reviewer does not know the original behaviour either, so a large generated diff is unverifiable. Agents are far more useful suggesting fixture cases you overlooked, while the local harness and snapshot remain the authority.
Avoid building fixtures from live customer data, and take care with security-sensitive parsers, where recording current behaviour can mean preserving a vulnerability. Characterization tests keep bugs by design, which is useful during a refactor and risky if forgotten afterwards.
Ready to take the first step towards unlocking opportunities, realizing goals, and embracing innovation? We're here and eager to connect.
11th Floor, O-Hub, Chandaka Industrial Estate, Infocity, Bhubaneswar, Odisha 751024
Level 4, 11 York Street Sydney Startup Hub Sydney, NSW – 2000
30 N. Đinh Nghệ, Phước Mỹ Sơn Trà, Đà Nẵng / Da Nang City – 550000
Level 25, AIDP Business Tower, Dubai Marina, United Arab Emirates
50 Beauchamp Street, Wellington, WGN 5028, New Zealand