Retry verification
Test retries in AI-generated code: detect duplicate writes
Run a local lost-reply simulation, observe a duplicate write, and test same-key replay plus conflicting input. Learn what the in-memory check cannot prove.
8 min read · Updated October 4, 2026
Put the lost reply after the write
Test a retry after the first operation completes but its response never reaches the caller. A test that throws before any side effect cannot tell you whether a second attempt duplicates a completed operation.
This guide uses a synthetic local note-creation fixture. You will observe two records without deduplication, then one record when the same operation key returns its saved result. The example uses Node.js 22 built-ins; no database, production endpoint, package install, or account is needed.
The question is different from whether the response has the right JSON shape: what happens to stored state when the caller cannot tell whether the first attempt succeeded?

State the retry contract
For this exercise, a logical operation has a key and a label. Retrying the same key with the same label must return the first record and leave exactly one record. Reusing that key with a different label must be rejected. Different keys represent different operations even when their labels match.
The caller must keep the key across attempts. Generating a new key on every retry would describe a new operation, so the receiver could not recognize the replay. A retry delay does not supply this identity.
These are deliberately chosen application semantics. A network timeout alone cannot establish that a write failed. Before automatically retrying a real operation, establish its documented replay behavior.
Run the complete local example
Save this as retry.test.mjs and run node --test retry.test.mjs:
import test from 'node:test';
import assert from 'node:assert/strict';
function writer(deduplicate) {
const records = [];
const completed = new Map();
return {
records,
write(key, label) {
if (deduplicate && completed.has(key)) {
const saved = completed.get(key);
if (saved.label !== label) throw new Error('key reused with different input');
return saved;
}
const saved = { id: records.length + 1, label };
records.push(saved);
if (deduplicate) completed.set(key, saved);
return saved;
},
};
}
function sendWithOneLostReply(store, key, label) {
let loseReply = true;
return () => {
const result = store.write(key, label);
if (loseReply) {
loseReply = false;
throw new Error('reply lost after write');
}
return result;
};
}
test('an unkeyed retry duplicates a completed write', () => {
const store = writer(false);
const send = sendWithOneLostReply(store, 'op-7', 'Draft');
assert.throws(send, /reply lost after write/);
assert.equal(store.records.length, 1);
const result = send();
assert.equal(store.records.length, 2);
assert.equal(result.id, 2);
});
test('the same key returns the original result after a lost reply', () => {
const store = writer(true);
const send = sendWithOneLostReply(store, 'op-7', 'Draft');
assert.throws(send, /reply lost after write/);
const result = send();
assert.deepEqual(result, { id: 1, label: 'Draft' });
assert.equal(store.records.length, 1);
});
test('a key cannot silently represent different input', () => {
const store = writer(true);
store.write('op-7', 'Draft');
assert.throws(() => store.write('op-7', 'Changed'), /different input/);
assert.equal(store.records.length, 1);
});
All three tests pass. The first intentionally demonstrates unsafe behavior: one record exists when the reply is lost, and the retry adds a second record with ID 2. A passing demonstration of a defect is not approval to ship that behavior.
The second checks the desired replay contract: the retry returns ID 1 and the record count stays one. The third proves that the same key cannot silently stand for changed input. Inspect the state and result together; a successful return alone would miss the duplicate.
Challenge the assertion, then restore
Temporarily change writer(true) to writer(false) in the second test only. Rerun it: that test must fail because the returned ID becomes 2. Restore true and rerun; all three tests must pass again.
This small check establishes that the assertion detects a missing replay guard. It does not establish that the guard survives a process restart or a race between servers. The in-memory map makes this example deterministic; it is not a durable storage design.
When examining an agent-generated retry loop, locate both the side effect and the retry trigger. Confirm the failure injection sits between the completed write and the observed reply, rather than before the operation starts.
Expand the test at the storage boundary
In a real repository, run through the actual operation handler and its durable storage layer. Keep any failure injection tightly scoped to the response boundary. Derive expected behavior from that operation's contract, rather than assuming every endpoint supports replay.
- Two overlapping requests using the same key must not both create records.
- A process restart after the write must not erase replay identity if the contract promises restart-safe retries.
- The saved result and side effect must not disagree after a partial failure.
- Separate logical operations must still succeed, including equal payloads with different keys.
- Key scope, input comparison, retention, and expired-key behavior need explicit rules.
Use separate regression cases for those promises. The fixture compares one label string; a real payload may need a defined comparison or fingerprint. Do not infer an exactly-once distributed guarantee from three sequential local tests.
Keep HTTP and local evidence separate
This example simulates the lost acknowledgment by throwing after the write. It does not send HTTP traffic or reproduce packet loss. Its evidence is narrow: the chosen sequential replay behavior and the recorded side-effect count.
RFC 9110's idempotent-method guidance explains why automatic retries depend on known semantics; the spelling of a retry helper cannot provide that guarantee. The Node.js 22 test runner documentation covers the local runner. Sources checked October 4, 2026.
For a review note, record the replay contract, where the reply is lost, the returned identity, the side-effect count, and the concurrency or durability cases still untested.
Make replay reasoning a manual rep
Before asking an agent to fix the loop, trace one ambiguous outcome yourself: the write happened, but the caller saw an error. Write down what the next attempt is allowed to change, then implement an independent assertion.
ManualMode's free review exercise uses synthetic patches and does not test your retry implementation. Three Gym reps and one Project rep are free after signup; a local repository regression can provide evidence for a bounded Project task, with raw project source staying local by default.
Start with evidence
Calibrate with three Gym reps, then verify one real Project task.
3 Gym + 1 Project reps free. Create an account; no card or public review required.