Skip to content
ManualMode guides

Refactor review

Characterization tests for AI refactoring: catch hidden mutation

Compare an AI refactor with existing behavior, catch an in-place sort that changes the caller's array, and verify a minimal correction with Node.js tests.

8 min read · Updated October 4, 2026

Compare observable behavior beyond the return value

For a behavior-preserving refactor, assert the promised effects on caller-owned inputs as well as the returned result. A differential test can compare old and new outputs while missing that the new implementation changes an array the caller still uses.

This guide demonstrates one synthetic AI-assisted refactor review. You will run a green output-comparison test beside a failing input-preservation test, then correct the refactor. Use Node.js 22 in a temporary directory; there is no package install, production access, or signup requirement.

A claim of unchanged behavior must include the state a caller observes after the call. The returned list is only one part of that evidence.

ManualMode diagram comparing a copy-then-sort refactor that preserves the original Lin and Ada array with an in-place sort that changes it, even though both outputs match.
Matching returned values can hide a changed caller-owned input.

Decide what is supposed to stay the same

Our small existing function accepts an array of names, returns a sorted copy, and leaves the input array untouched. For the two plain ASCII strings in this fixture, the sorted result is Ada followed by Lin. Keeping caller state unchanged is part of the chosen contract.

The agent's proposed simplification removes the copy and returns names.sort(). The return value still looks right. But JavaScript's sort method changes the array it receives, so a caller can observe different state after the call.

A characterization test records current behavior. Before treating that behavior as a compatibility requirement, compare it with documentation and actual consumers. An old bug should not become correct merely because a new test preserves it.

Run the complete comparison

Save this as refactor.test.mjs and run node --test refactor.test.mjs:

import test from 'node:test';
import assert from 'node:assert/strict';

function legacy(names) { return [...names].sort(); }
function refactored(names) { return names.sort(); }

test('output matches the legacy implementation', () => {
  assert.deepEqual(refactored(['Lin', 'Ada']), legacy(['Lin', 'Ada']));
});

test('the refactor preserves the caller array', () => {
  const input = ['Lin', 'Ada'];
  const before = [...input];
  assert.deepEqual(refactored(input), ['Ada', 'Lin']);
  assert.deepEqual(input, before);
});

The output comparison passes: both functions produce Ada followed by Lin. The second test fails at the input assertion. Its expected input remains Lin followed by Ada; the actual input has become Ada followed by Lin. The observed total is one passing test and one failing test.

Notice the distinct arrays in the first test. Reusing the same mutable fixture across both implementations could let the first call change the input that the second receives, making the comparison harder to interpret.

The second expectation is independent of the refactored algorithm: it snapshots the caller's original values and checks that the function leaves them alone. The expected output is also explicit for this known fixture.

Restore the copy and verify both promises

Change only the refactored function to function refactored(names) { return [...names].sort(); }. Run the same command again. Both tests pass: the returned list is sorted and the original list keeps its order.

This copy is shallow. It creates another array containing the same element references; it does not clone nested objects. Our fixture uses strings and checks array order, so that limitation does not affect the demonstrated result. A real function that modifies object fields would need additional assertions.

Do not replace the failing input expectation with the new sorted order just to make the suite green. If changing caller state is intentional, it is a behavior change requiring a revised contract and consumer review, rather than a behavior-preserving refactor.

Build a small compatibility matrix

Before accepting a broader cleanup, identify the observable promises it could affect. Use focused fixtures rather than a huge snapshot of unrelated output.

  • Returned values and permitted output order.
  • Caller-owned arrays, objects, or other state that must remain unchanged.
  • Which failures are thrown or returned for documented invalid inputs.
  • External calls, writes, or emitted events whose count and order matter.
  • References whose identity is explicitly part of the API contract.

Do not automatically assert every implementation detail. A changed allocation pattern or private helper name may be irrelevant. Choose properties that callers or documented requirements can observe. Give the old and new implementations isolated fixtures so one cannot contaminate the other.

For each affected promise, ask whether the old behavior is required, tolerated, or known to be wrong. A bug fix can legitimately change the answer; document it separately from a pure cleanup.

Keep the evidence narrower than compatibility

Two tests do not establish equivalence for every input. This example does not cover locale-aware ordering, nonstring inputs, large arrays, performance, nested mutation, persistence, or async effects. The existing implementation could also have undiscovered defects that an output comparison would repeat.

Use characterization to locate unexpected differences, then use independent contract assertions to decide which behavior is right. In a real repository, rerun the affected caller tests and broader checks after the focused case passes.

MDN's Array.sort reference documents its in-place behavior. The Node.js 22 test runner documentation covers the runner used here. Sources checked October 4, 2026.

Martin Fowler describes refactoring as changing structure while preserving observable behavior. Here, the caller's array is one observable part of that contract.

Reserve the judgment for a manual rep

Before accepting an agent's cleanup, read one affected caller and write down its state before and after the function runs. Then build one assertion without copying the new implementation into the expected value.

ManualMode's free review exercise uses synthetic patches; it does not analyze your refactor or run this test. Three Gym reps and one Project rep are free after signup. For a real repository task, keep the code local and use a bounded regression as evidence of your manual work.

Start with evidence

Calibrate with three Gym reps, then verify one real Project task.

3 Gym + 1 Project reps free. Create an account; no card or public review required.

Start free