Software Engineering
AI code review vs human code review: what each one is good at
A practical comparison of automated AI review and human review: what AI catches, what it misses, how to combine them, and a setup that keeps review honest.
By Raktim Ranjit · Published · 3 min read
Short answer: use both, for different jobs. AI review is fast, tireless and good at surface problems: missing null checks, unused code, inconsistent naming, obvious security slips and gaps in tests. Human review is better at whether the change should exist, whether it fits the design, and what it will do to the people who maintain it. Let the tool do the first pass so people spend their attention on the second.
What does AI code review catch well?
- Typos, dead code, unused imports and variables.
- Missing error handling and unchecked return values.
- Common security mistakes like string-built SQL, missing authorisation checks on a new route, and secrets in a diff.
- Inconsistent style and naming when it has seen the rest of the repository.
- Missing tests for a new branch.
- Mismatches between a function's comment and its behaviour.
It does this on every pull request, at any hour, without getting bored on the 40th file.
What does it miss?
- Intent. It reads the diff, not the conversation that led to it. It cannot know that this feature was cancelled last week.
- System-wide effects. A change that is correct locally and breaks a job running in another service.
- Design fit. Whether this should be a new table or a column, a new service or a function.
- Operational reality. Migrations that lock a large table, rollout order, and what happens during a deploy.
- Taste for the codebase. Whether the team wants this pattern spreading.
It also produces false positives and confident wrong claims. A comment that is wrong costs the author time to disprove.
What does a human reviewer add?
Context and accountability. A person can ask why, push back on scope, notice that two changes should be separate, and teach. Review is also how knowledge moves around a team. If an assistant approves everything and nobody reads the code, the team loses its understanding of its own system even while the tests are green.
How should you combine them?
- Run automated review on every pull request as a first pass.
- Let the author resolve the easy comments before a person looks.
- Ask humans to focus on design, behaviour, risk and tests, and tell them they do not need to repeat what the tool said.
- Require a human approval to merge. Never let a bot be the only approver for production code.
- Keep the diff small. Both humans and models review small diffs far better.
Does AI-written code need stricter review?
Yes, in a specific way. Generated code tends to look right. It is consistent, well formatted and commented, which lowers a reviewer's guard. Read it as you would read code from a new contractor you have not worked with yet. Check that the tests would fail if the code were wrong, that dependencies it added exist and are the ones you meant, and that it did not quietly change behaviour outside the request.
How do you measure whether it helps?
Track what you can count: time from pull request opened to first feedback, time to merge, how many automated comments the author accepts, and how many bugs reach production that review should have caught. If the acceptance rate of bot comments is under about a third, the tool is producing noise and people will start ignoring it, which is worse than having none.
A review checklist that works for both
- Does the change do what the description says, and nothing else?
- What happens on bad input, an empty result and a timeout?
- Who is allowed to call this, and is that checked?
- How will we know it is broken in production?
- Can we roll it back?
Those five questions are cheap for a person to ask and hard for a tool to answer without context. They are where a human still earns their place.
Author
Raktim Ranjit is a software engineer and the founder of NodeDR Infotech. He builds and maintains the software described here.