How to Review AI-Generated Code: A Practical Checklist
AI-assisted code should meet the same engineering bar as human-written code. The difference is where reviewers should spend attention: context mismatches, plausible-but-wrong assumptions, and tests that can accidentally validate the same mistake as the implementation.
Start with intent, not style
Before reading the implementation line by line, compare the diff with the issue, acceptance criteria, and explicit non-goals. An agent can produce polished code for a subtly different problem. A review that starts with naming and formatting can miss that larger mismatch.
- What user-visible or system behavior was supposed to change?
- Which files and layers genuinely needed to change?
- Did the implementation introduce unrelated refactors, dependencies, generated churn, or configuration changes?
Seven failure modes worth checking
1. It solved a plausible problem, not the requested problem
Trace each material change back to a requirement. If you cannot explain why a branch, field, migration, or side effect exists, ask for evidence before treating it as necessary.
2. It changed more than necessary
Separate required work from opportunistic cleanup. Extra surface area makes regressions harder to isolate and rollback harder to reason about.
3. The tests repeat the same mistaken assumption
Look for at least one test that proves observable behavior independently of implementation details. A test generated from the same interpretation as the code can be internally consistent and still wrong.
4. Failure, timeout, or retry semantics were guessed
For external calls and background work, verify timeout meaning, idempotency, retry boundaries, duplicate side effects, partial success, and what the caller sees when dependencies fail.
5. Authorization or ownership checks moved accidentally
Trace who can read, create, update, approve, or delete affected resources. Pay special attention when logic moved between UI, API, service, job, or database layers.
6. Migration and rollout ordering were ignored
Check whether old and new application versions can coexist during deployment, whether existing rows remain valid, and whether rollback is still possible after the schema or data changes.
7. The reviewer invents findings because it was asked to find issues
A useful review can end with no blocking issues found. Require a concrete code, test, behavior, security, data, or operational basis for blockers; keep questions and optional improvements separate.
A fast evidence-first review sequence
- Intent: confirm the requested behavior and non-goals.
- Correctness: exercise happy path, boundaries, empty states, failures, and concurrency where relevant.
- Security and privacy: verify authorization, ownership, secrets, sensitive data, and trust boundaries.
- Data: inspect migrations, compatibility, destructive operations, retries, and recovery.
- Tests: require coverage for materially changed behavior and at least one meaningful failure path.
- Delivery: understand deploy order, observability, rollback, and operational failure modes.
- Maintainability: raise it only when it creates credible defect or operational risk.
Use the checklist on a real PR
For a compact version you can keep beside a pull request, use the pull request review checklist. If you want an AI reviewer to follow the same evidence standard, use the AI code review prompt.