Testing the tests
Mutation testing AI-generated tests: prove they catch a wrong answer
Check AI-generated tests with a runnable Node.js boundary example: make one deliberate bug survive, add an independent assertion, and verify it fails.
8 min read · Updated October 2026
A green generated suite still needs a challenge
Change one piece of correct behavior on purpose, then ask whether the tests notice. This is a small manual mutation check. It helps you assess assertions when a coding agent wrote both the implementation and its tests.
The useful question is specific: would these tests reject a plausible wrong answer? Passing tests alone cannot answer it. You need a stated contract, one controlled change, and an observed assertion failure.

This exercise is synthetic, runs locally, and uses no ManualMode customer code. You need Node.js 22 and two files in a temporary directory; no package install or account is required.
Write an oracle outside the implementation
Our invented requirement is: valid subtotals are nonnegative integer cents; shipping costs 500 cents below 5000 cents and is free at 5000 cents or more. Using integer cents avoids making floating-point rounding part of this exercise.
Save this implementation as shipping.mjs:
// shipping.mjs
// Contract: nonnegative integer cents, free shipping at 5000 or more.
export function shippingFee(subtotalCents) {
return subtotalCents >= 5000 ? 0 : 500;
}Save these two tests beside it as shipping.test.mjs:
// shipping.test.mjs
import test from 'node:test';
import assert from 'node:assert/strict';
import { shippingFee } from './shipping.mjs';
test('below the threshold costs 500 cents', () => {
assert.equal(shippingFee(4999), 500);
});
test('above the threshold has free shipping', () => {
assert.equal(shippingFee(5001), 0);
});Run node --test shipping.test.mjs. Both tests pass. Each expected value comes directly from the written shipping rule. An expectation such as subtotal >= 5000 ? 0 : 500 would repeat the algorithm and could carry the same mistake.
The Node.js 22 test runner documentation describes the runner used here. This example uses its standard test function and strict equality assertions.
Make an off-by-one bug survive
In shipping.mjs, replace only the return line with:
// Temporary mutation in shipping.mjs: >= becomes >.
return subtotalCents > 5000 ? 0 : 500;Run the same command again. Both tests still pass. The mutated function now charges shipping at exactly 5000 cents, which violates the contract, but neither assertion uses that input.
You have found a surviving mutation: an intentional behavior change the selected suite did not reject. The evidence is the changed operator and the two passing tests, rather than a guess that the suite probably misses edge cases.
Restore the original >= operator before proceeding. Keep deliberate bugs in a disposable copy or a small local diff; never ship one as part of this exercise.
Add the missing assertion and observe red, then green
Append this test to shipping.test.mjs:
test('exactly 5000 cents has free shipping', () => {
assert.equal(shippingFee(5000), 0);
});- With the original
>=implementation, run the suite: three tests pass. - Change
>=to>again: two tests pass and the new test fails, receiving500instead of0. - Restore
>=and rerun: three tests pass. Check the diff to confirm the mutation is gone.
The failed equality assertion matters. An import error, syntax error, or unavailable service would not establish that the test understands the threshold. The implementation must load and produce the wrong observable result.
Now try returning 0 unconditionally. The below-threshold test should fail. Restore the implementation afterward. This second check challenges whether the suite distinguishes the paid-shipping path at all.
Choose mutations from the real change's risks
For an AI-assisted patch in your repository, pick one changed decision that could plausibly be wrong. A comparison boundary is a good first target; other candidates depend on the actual contract.
- Replace one returned value with a constant and check that callers notice.
- Remove an ownership condition in a local test copy and verify a different owner's fixture is rejected by an independent assertion.
- Reverse an ordering decision and check the documented output order.
Change one thing at a time. Run the focused suite first, record the assertion, then restore and run the repository's normal checks. If a mutation survives, decide whether it exposes a missing assertion or preserves the specified behavior for every valid input.
For larger suites, automated tools can generate and run many such changes. The Stryker mutation-testing introduction explains detected and surviving mutations. This walkthrough performs two manual checks; it does not run Stryker or calculate a mutation score.
Keep the result narrower than the claim
These tests distinguish the threshold operator and a constant-free-shipping mistake. They do not validate negative, fractional, nonnumeric, or enormous inputs; the example declares those outside its input contract. A real checkout also needs its own currency, tax, discount, persistence, and integration rules.
Detecting a few mutations cannot prove that a requirement is correct, that every generated test is useful, or that a patch is secure. A surviving change can be equivalent under the contract; repeatedly adding assertions to inflate a score can waste time.
Leave a review note another engineer can reproduce: requirement, mutated line, command, failing assertion, restored diff, and remaining unverified behavior. Use the AI-generated code testing workflow for the broader merge decision and the review checklist for reading the patch.
Turn one test gap into a manual coding rep
Before asking an agent for another test, write the missing assertion yourself and predict the failure. That small act practices reading a contract, choosing an input, and checking an observable result.
ManualMode offers tailored Gym practice and bounded Project tasks in your own repository. In Project Mode, a connected agent can propose a task; ManualMode approves and reserves its scope before the manual rep counts. Raw project source stays local by default. This article does not inspect your tests or upload your patch.
Try the public review exercise on synthetic patches, or read how to practice in your own projects. The complete mutation exercise above is available without signup.
Start with evidence
Calibrate with three Gym reps, then verify one real Project task.
3 Gym + 1 Project reps free. Create an account; no card or public review required.