How to Review AI-Generated Code: A Practical Checklist

AI-assisted code should meet the same engineering bar as human-written code. The difference is where reviewers should spend attention: context mismatches, plausible-but-wrong assumptions, and tests that can accidentally validate the same mistake as the implementation.

Start with intent, not style

Before reading the implementation line by line, compare the diff with the issue, acceptance criteria, and explicit non-goals. An agent can produce polished code for a subtly different problem. A review that starts with naming and formatting can miss that larger mismatch.

Seven failure modes worth checking

1. It solved a plausible problem, not the requested problem

Trace each material change back to a requirement. If you cannot explain why a branch, field, migration, or side effect exists, ask for evidence before treating it as necessary.

2. It changed more than necessary

Separate required work from opportunistic cleanup. Extra surface area makes regressions harder to isolate and rollback harder to reason about.

3. The tests repeat the same mistaken assumption

Look for at least one test that proves observable behavior independently of implementation details. A test generated from the same interpretation as the code can be internally consistent and still wrong.

4. Failure, timeout, or retry semantics were guessed

For external calls and background work, verify timeout meaning, idempotency, retry boundaries, duplicate side effects, partial success, and what the caller sees when dependencies fail.

5. Authorization or ownership checks moved accidentally

Trace who can read, create, update, approve, or delete affected resources. Pay special attention when logic moved between UI, API, service, job, or database layers.

6. Migration and rollout ordering were ignored

Check whether old and new application versions can coexist during deployment, whether existing rows remain valid, and whether rollback is still possible after the schema or data changes.

7. The reviewer invents findings because it was asked to find issues

A useful review can end with no blocking issues found. Require a concrete code, test, behavior, security, data, or operational basis for blockers; keep questions and optional improvements separate.

A fast evidence-first review sequence

  1. Intent: confirm the requested behavior and non-goals.
  2. Correctness: exercise happy path, boundaries, empty states, failures, and concurrency where relevant.
  3. Security and privacy: verify authorization, ownership, secrets, sensitive data, and trust boundaries.
  4. Data: inspect migrations, compatibility, destructive operations, retries, and recovery.
  5. Tests: require coverage for materially changed behavior and at least one meaningful failure path.
  6. Delivery: understand deploy order, observability, rollback, and operational failure modes.
  7. Maintainability: raise it only when it creates credible defect or operational risk.
Decision rule: request changes for a credible ship risk supported by evidence. Ask for evidence when a risk cannot yet be verified. Otherwise approve instead of manufacturing comments.

Use the checklist on a real PR

For a compact version you can keep beside a pull request, use the pull request review checklist. If you want an AI reviewer to follow the same evidence standard, use the AI code review prompt.

Get the free Markdown checklist

See a fictional sample audit