Skip to content

A checklist for reviewing AI-generated code

By SunnyKumar Jonwal 10 min read

Code from an AI assistant tends to fail in a specific way: it looks right. The formatting is clean, the names are sensible, the comments are tidy, and it passes the happy path. That polish lowers your guard, which is exactly when subtle mistakes slip through.

Reviewing it isn't fundamentally different from reviewing a colleague's pull request, but the failure modes have a different shape. A human junior developer makes typos and misunderstands requirements. A model makes plausible-looking decisions that were never grounded in your system. This post is a working checklist for catching them.

Start with a stance

Assume the code is a draft from a fast, well-read contributor who didn't attend your planning meeting, can't see production, and has never been paged at 3 a.m. It knows patterns. It doesn't know your users, your data, or your incident history.

That framing keeps two errors at bay: trusting it too much because it sounds authoritative, and distrusting it so much that you rewrite everything. Both waste time. Review for the places where a confident pattern-matcher goes wrong.

First pass: did it do what you asked, and only that?

Before reading line by line, look at the shape of the change. How many files were touched? Does that match the size of the request? Agents are eager, and they'll often "improve" things nearby: renaming variables, reformatting a file, upgrading a dependency, refactoring a function you didn't mention.

Scope creep isn't harmless. Every unrequested change is something you now have to review and something that could break. If a diff for a small bug fix touches fifteen files, stop and ask why. Either the fix is bigger than you thought, or the agent wandered, and it's worth knowing which.

Check what got deleted as well as what got added. Removed tests, removed error handling, removed validation, and removed comments explaining an odd decision all deserve attention.

Read the tests first

If the change includes tests, read them before the implementation. They reveal what the author thought the behavior should be, and they're often where the trouble hides.

Watch for tests that don't test anything: assertions that always pass, mocks so extensive that the test only verifies the mock, expected values copied from whatever the code currently returns. A test written to match the implementation can't catch an implementation bug.

Watch for tests that were weakened to get green. If a failing test was edited so that it passes, or a skipped marker was added, or an assertion loosened from an exact value to "not null," that's a red flag. Ask what the original test was protecting.

Then check coverage of the awkward cases: empty input, huge input, duplicates, null values, boundary values, concurrent access, failure of a dependency. Models write the happy path readily and the unhappy paths reluctantly.

Correctness: where plausible code goes wrong

Now the implementation. Some recurring trouble spots.

Assumptions about data. The code assumes a field is always present, a list is never empty, an ID is unique, a timestamp is in UTC. Check each against reality. The model guessed, and you know the real data.

Off-by-one and boundary errors. Pagination, date ranges, slicing, loops with inclusive and exclusive bounds. These are classic and easy to skim past.

Time and money. Time zones, daylight saving, floating-point arithmetic on currency, rounding rules. If the code handles either, slow down.

Error handling. Look for broad catch blocks that swallow exceptions, empty handlers, and fallbacks that hide failures by returning a default. A silent failure is worse than a loud one.

Concurrency. Read-modify-write sequences without locks or transactions, race conditions between check and use, jobs that aren't safe to run twice. Models often write code that's correct for one request at a time.

API misuse. Functions that don't exist in your version of a library, deprecated calls, or parameters in the wrong order. The code may be idiomatic for a different major version than the one you run.

Invented behavior. It calls a method that seems like it should exist. In typed languages the compiler catches this. In dynamic ones, you might only find out at runtime.

Running the code is the best defense against all of these. Don't just read it. Execute it with realistic data, including a few nasty inputs, and watch what happens.

Security: the checks that matter most

Security bugs are where AI-written code costs the most, because they're invisible in normal use. A short list to run every time.

Authorization. Does every endpoint check that the current user is allowed to touch this specific record, not just that they're logged in? Missing object-level checks are among the most common findings in real audits, and generated code omits them readily. Here's a typical shape in Laravel:

// Looks fine. Any logged-in user can read any invoice by guessing ids.
public function show(Invoice $invoice)
{
    return new InvoiceResource($invoice);
}

// Scoped: a policy decides whether this user may view this invoice.
public function show(Invoice $invoice)
{
    $this->authorize('view', $invoice);

    return new InvoiceResource($invoice);
}

Input handling. Is every input validated? Is user input ever concatenated into SQL, shell commands, file paths, or HTML? Parameterized queries, escaped output, and validated paths should be the norm. Watch for raw queries and exec-style calls.

Secrets. No keys, tokens, or passwords in code, tests, or logs. Check example configs and sample values too, since generated code sometimes includes realistic-looking credentials.

Mass assignment and over-exposure. Are models filled from raw request data? Do responses include fields the caller shouldn't see, such as internal flags or other users' data?

Unsafe defaults. Disabled TLS verification "for testing," permissive CORS, debug mode enabled, overly broad file permissions. These often appear as quick fixes to make something work.

Logging. Does it log passwords, tokens, or personal data?

The general principle from least privilege for AI agents applies to code too: anything that grants access deserves a second look.

Dependencies: check that they exist

This one is specific to AI-written code. Models sometimes suggest package names that don't exist. Attackers know this, and there have been reports and research showing that they can register those names with malicious code in wait, a technique sometimes called slopsquatting. Installing a hallucinated dependency blindly is a supply-chain risk.

Before accepting any new dependency, confirm the package exists on the official registry under exactly that name, check who publishes it, how long it has existed, how many people use it, and whether it's maintained. Look at what it actually does. Prefer the standard library or something already in your project over a new package for small jobs. Pin versions, and let your normal dependency scanning run on the change.

Also check that the version being used matches the API the code calls. An import that works in version 2 of a library may not exist in version 4.

Performance traps

Generated code often works on ten rows and falls over on ten million. Some things to look for:

  • Queries inside loops. The classic N+1 pattern, where you fetch a list and then query each item separately.
  • Unbounded reads: loading a whole table into memory, or an entire file at once.
  • Missing indexes for new query patterns.
  • Repeated work inside hot paths, like recomputing something on every iteration.
  • Synchronous calls to slow services where a queue would belong.

An example in Laravel:

// N+1: one query for invoices, then one per invoice for its customer.
$invoices = Invoice::latest()->take(50)->get();
foreach ($invoices as $invoice) {
    echo $invoice->customer->name;
}

// Eager loading: two queries total.
$invoices = Invoice::with('customer')->latest()->take(50)->get();

You don't need to optimize prematurely, but you should know when a change will turn one query into a thousand.

Maintainability

Even correct code can be a burden. Check for duplication of logic that already exists elsewhere in the codebase, since models don't always know your helpers. Look for over-engineering: abstractions, interfaces, and configuration options nobody asked for. Check naming against your conventions, dead code and leftover debug statements, and comments that restate the code without explaining why. If a future teammate would have to ask "why is this here?", that's a comment or a simplification waiting to happen.

Most important: can you explain how it works? If you can't, don't merge it. Being the person who ships code nobody understands is how incidents start.

Process habits that help

A few habits make review both faster and more reliable.

Keep changes small. A 40-line diff can be reviewed carefully, and a 900-line diff gets skimmed. Ask the agent to work in small steps and commit each one, as in the Claude Code workflow guide.

Use the tools. Run linters, type checkers, static analysis, and the whole test suite, since the agent may have run only a subset. Automated checks are tireless, and they catch the boring stuff so your attention can go to the subtle stuff.

Ask for an explanation, and read it critically. "Why did you choose this approach? What did you consider and reject? What could go wrong?" Explanations aren't proof, but they often expose an unchecked assumption.

Consider a second opinion from another model or a fresh session, pointed at the diff with a specific brief such as "review for authorization gaps and injection." It's useful as an extra pass. It's not a substitute for your own read, because reviewers built from the same kind of model can share blind spots.

Write the missing test yourself. If you're unsure whether the behavior is right, writing a test that pins it down is the fastest way to find out.

A mini-review, start to finish

To see how this plays out, imagine a request: "Add an endpoint that lets a customer download their invoice as a PDF." The agent returns a diff. Here's how a review might go.

You glance at the file list first. Eight files changed for a single endpoint, which is more than expected. Two are the route and controller, one is a new PDF service, one is a view template, and one is the test. The other three are a config file, a composer lockfile, and an unrelated model that got reformatted. You ask about the last three. The config change adds a new PDF library, the lockfile follows from that, and the model edit is just whitespace. You'll ask for the whitespace change to be reverted, and you'll look at the library.

The library is one you haven't used, so you check its registry page: it exists, it's maintained, and it has a healthy user base. Fine. But the project already includes another PDF package, so you ask why a second was added. The agent hadn't noticed the first one. That's a duplicate dependency avoided.

Next, the test. It creates an invoice, requests the PDF as the owner, and asserts a 200 response. There's no test for a different user's invoice, and the controller loads the invoice by ID from the route without any policy check. That's the missing authorization from earlier: any logged-in customer can download anyone's invoice by changing the number in the URL. You ask for a policy check and a test proving another user gets a 403.

Finally you run it. The PDF generates, and with a 3,000-line invoice, it takes twelve seconds and a lot of memory, because the whole invoice is rendered synchronously. You decide that's acceptable for now and note it as follow-up work.

That review took about ten minutes and found three real problems: a duplicate dependency, an authorization hole, and a scope leak. None was visible from the happy-path test.

Red flags at a glance

Some signs that deserve an immediate second look: tests changed alongside the code they test, catch-all exception handlers, new dependencies you haven't heard of, hardcoded credentials or URLs, disabled security features "for now," queries or shell commands built from strings, unexplained changes in unrelated files, and confident comments that describe behavior the code doesn't have.

The point of all this

None of this argues against using AI for code. It argues for treating the output as a proposal that you're accountable for. The teams getting real value from these tools tend to pair them with strong tests, small changes, and reviewers who know what to look for. With that in place, the speed is real and the risks are manageable.