给代码审查者记功,解决 AI 生成代码的审查走过场问题
Fix Your AI Slop Problem by Giving Reviewers Credit
一位开发者发现,AI 写的代码经 AI 审查、人类再把 PR 粘进 AI 审查后盖章通过,连明显 bug 都没被发现。他建议把 review 作为子任务加到 sprint board 上,让审查工作量可见并获得认可,从而避免审查者沦为 AI 审查的传声筒。
I was reviewing a PR recently. See if you can spot the bug:
Obviously I don’t work at a library — I swapped out the real calls. But it was truly that simple. The # Branch C doesn’t have a printer isn’t added here for context either, it’s there, in the code. Any testing we'd do would pass, but the moment we deployed to branch C, we'd start seeing failures.
Maybe a better codebase wouldn’t have made this type of mistake possible. But hey, don’t blame me — production code bases are rarely perfect, and as far as ones I’ve seen go, this one isn’t too bad.
Still, good old blind copy-and-pasting would’ve prevented this. Scrolling up should’ve prevented this.
Instead, an AI coded it up, an AI model reviewed it, a human reviewer pasted it into an AI to review it, and then stamped it.
I was tagged as the second (human) reviewer. This time I happened to read the code, but I’ll admit — most times, it would’ve just pasted the PR into AI and missed the same obvious bug.
I’ve always been taught that a reviewer owns the code as much as the author. And I’d always try and check a bunch of things: does it work, does it fit, will anyone understand it in a year. But there was one thing I’d never had to check before — whether the author ever read it themselves.
That came guaranteed. You couldn’t have typed staff.print_document() without hitting ctrl+f and finding an example of how it was used. Yet, today, people can ship code they’ve never seen before. The reviewer could plausibly be the first human to actually read the code.
The bug I saw didn’t require any special knowledge to catch. The fact AI even missed it was honestly pretty shocking to me. But at the same time, having a circle of AI’s read the same code, with near identical prompts, often running the same model1, and expecting to discover something different makes no sense.
We were mistaking a second (and third and fourth) AI pass for a human review just because it was a human copy and pasting the PR into their favourite AI harness.
At what point should you just cut the human out of the loop and save everyone some time?
I’ve talked with a few teams at various companies over the last year, and more than one has ditched the human entirely. A few would have the AI make a judgment call — is this change risky enough to warrant a human reviewer?
Personally, I’m still a fan of having human reviewers for the vast majority of work. Maybe I’m behind the curve.
Either way, this isn’t a blog post about deciding how much human you want. It’s that once you decide that number, whether it’s 0%, 100%, or something in the middle, then you should then take that direction with intention.
Today, many of us are living in no man’s land. We still have the human reviewer requirement from the past — but the human is just a proxy for yet another AI review.
At first glance a tempting fix might be to ban AI from reviews until the reviewer has at least read it once themselves. But you can't police what someone does in their own terminal, and even if you could, it’s probably not addressing the true root cause.
There are few things more infuriating than reviewing a PR or a doc where you know you’re about to spend more time reading it than the author spent writing it. They spent 5 minutes prompting AI, just for you to spend an hour reviewing it. At what point are you the one building the feature, by proxy?
And these hours that you spend reviewing code — you never get any credit. Pasting PRs into AI isn’t laziness. It’s the logical decision if you prioritize your own time.
In my experience, it’s rare for people to want to game the system. Most people want to do good, honest work and be recognized for it. They just don’t want to do invisible, uncredited labour.
The solution to this problem evaded me for a few weeks, but I actually think it’s quite simple: add reviews as subtasks on your sprint board.
Pre-AI, managing that overhead would’ve been a pain in the ass. These days AI can do it for you — hook it into your PR pipeline, have it create the ticket automatically.
And I know, I know. It’s bad practice to add tasks mid-sprint. But fuck it. Times have changed. Unless you want to slow down shipping — and no one wants that — you’re going to have to drag those tasks into the current sprint.
A review ticket does two things. First, it forces you to estimate it. Sure — it’s probably a 1 pointer, but spending 30 seconds estimating it is a signal that it’s a real thing that we value.
Second, it makes the cost visible. When a 3-point feature spawns a 3-point review, that effort finally shows up side-by-side with the feature work itself. And for the first time ever, the reviewer will no longer be punished for actually doing their due diligence.
Now you would be right to point out — how is adding reviewers to sprint boards a solution? People can still use AI to save time.
The answer is that this is a one-two punch. Giving reviewers credit sets you up to build a team culture where you’re rewarded for spending time on reviews and finding issues early. Adding reviews to sprint boards is putting up one extra wall that has to be broken down for your culture to turn to AI slop.
So give people the recognition they deserve for reviewing PRs properly, and next time someone hits approve, it’ll actually mean something.
Even using different models might not matter much — one 2025 study had found that when two different models get a question wrong, they pick the same wrong answer 60% of the time.
No posts
来源:Hacker News · AI · danunparsed.com