Refactor review
Characterization tests for AI refactoring: catch hidden mutation
Compare an AI refactor with existing behavior, catch an in-place sort that changes the caller's array, and verify a minimal correction with Node.js tests.
8 min read · Updated October 4, 2026
Compare observable behavior beyond the return value
For a behavior-preserving refactor, assert the promised effects on caller-owned inputs as well as the returned result. A differential test can compare old and new outputs while missing that the new implementation changes an array the caller still uses.
This guide demonstrates one synthetic AI-assisted refactor review. You will run a green output-comparison test beside a failing input-preservation test, then correct the refactor. Use Node.js 22 in a temporary directory; there is no package install, production access, or signup requirement.
A claim of unchanged behavior must include the state a caller observes after the call. The returned list is only one part of that evidence.

Decide what is supposed to stay the same
Our small existing function accepts an array of names, returns a sorted copy, and leaves the input array untouched. For the two plain ASCII strings in this fixture, the sorted result is Ada followed by Lin. Keeping caller state unchanged is part of the chosen contract.
The agent's proposed simplification removes the copy and returns names.sort(). The return value still looks right. But JavaScript's sort method changes the array it receives, so a caller can observe different state after the call.
A characterization test records current behavior. Before treating that behavior as a compatibility requirement, compare it with documentation and actual consumers. An old bug should not become correct merely because a new test preserves it.
Run the complete comparison
Save this as refactor.test.mjs and run node --test refactor.test.mjs:
import test from 'node:test';
import assert from 'node:assert/strict';
function legacy(names) { return [...names].sort(); }
function refactored(names) { return names.sort(); }
test('output matches the legacy implementation', () => {
assert.deepEqual(refactored(['Lin', 'Ada']), legacy(['Lin', 'Ada']));
});
test('the refactor preserves the caller array', () => {
const input = ['Lin', 'Ada'];
const before = [...input];
assert.deepEqual(refactored(input), ['Ada', 'Lin']);
assert.deepEqual(input, before);
});
The output comparison passes: both functions produce Ada followed by Lin. The second test fails at the input assertion. Its expected input remains Lin followed by Ada; the actual input has become Ada followed by Lin. The observed total is one passing test and one failing test.
Notice the distinct arrays in the first test. Reusing the same mutable fixture across both implementations could let the first call change the input that the second receives, making the comparison harder to interpret.
The second expectation is independent of the refactored algorithm: it snapshots the caller's original values and checks that the function leaves them alone. The expected output is also explicit for this known fixture.
Restore the copy and verify both promises
Change only the refactored function to function refactored(names) { return [...names].sort(); }. Run the same command again. Both tests pass: the returned list is sorted and the original list keeps its order.
This copy is shallow. It creates another array containing the same element references; it does not clone nested objects. Our fixture uses strings and checks array order, so that limitation does not affect the demonstrated result. A real function that modifies object fields would need additional assertions.
Do not replace the failing input expectation with the new sorted order just to make the suite green. If changing caller state is intentional, it is a behavior change requiring a revised contract and consumer review, rather than a behavior-preserving refactor.
Build a small compatibility matrix
Before accepting a broader cleanup, identify the observable promises it could affect. Use focused fixtures rather than a huge snapshot of unrelated output.
- Returned values and permitted output order.
- Caller-owned arrays, objects, or other state that must remain unchanged.
- Which failures are thrown or returned for documented invalid inputs.
- External calls, writes, or emitted events whose count and order matter.
- References whose identity is explicitly part of the API contract.
Do not automatically assert every implementation detail. A changed allocation pattern or private helper name may be irrelevant. Choose properties that callers or documented requirements can observe. Give the old and new implementations isolated fixtures so one cannot contaminate the other.
For each affected promise, ask whether the old behavior is required, tolerated, or known to be wrong. A bug fix can legitimately change the answer; document it separately from a pure cleanup.
Keep the evidence narrower than compatibility
Two tests do not establish equivalence for every input. This example does not cover locale-aware ordering, nonstring inputs, large arrays, performance, nested mutation, persistence, or async effects. The existing implementation could also have undiscovered defects that an output comparison would repeat.
Use characterization to locate unexpected differences, then use independent contract assertions to decide which behavior is right. In a real repository, rerun the affected caller tests and broader checks after the focused case passes.
MDN's Array.sort reference documents its in-place behavior. The Node.js 22 test runner documentation covers the runner used here. Sources checked October 4, 2026.
Martin Fowler describes refactoring as changing structure while preserving observable behavior. Here, the caller's array is one observable part of that contract.
Reserve the judgment for a manual rep
Before accepting an agent's cleanup, read one affected caller and write down its state before and after the function runs. Then build one assertion without copying the new implementation into the expected value.
ManualMode's free review exercise uses synthetic patches; it does not analyze your refactor or run this test. Three Gym reps and one Project rep are free after signup. For a real repository task, keep the code local and use a bounded regression as evidence of your manual work.
Start with evidence
Calibrate with three Gym reps, then verify one real Project task.
3 Gym + 1 Project reps free. Create an account; no card or public review required.