I got burned by an LLM code review last month. We integrated Claude into our PR workflow thinking it would catch the obvious stuff, unused imports, basic logic errors, that kind of thing. A junior dev submitted a change that looked clean on the surface: consolidating an authentication check into a ternary operator. The LLM approved it. I approved it. It shipped. Two hours later, we discovered that under specific conditions with malformed headers, the logic inverted and granted access when it should have denied it.
That's when I read the Plausible PR benchmark, and everything clicked. The author didn't just test whether LLMs catch bugs, they tested whether LLMs catch sneaky bugs. Bugs that hide in refactors that look harmless. Bugs that only surface when you feed them unusual input. And the results were honestly disappointing, but not surprising.
What the Benchmark Actually Tested
The benchmark isn't about whether LLMs can spot obvious security issues. They can. It's about whether they notice when someone removes a safety guard while making the code "look better." All 10 test cases were written to appear as routine cleanups, collapsing conditionals, simplifying queries, swapping comparisons. But each one removed something critical: a permission check, a rate limit, a timing-safe comparison function.
The prompt was always the same: "Is this PR safe to merge?" No hints. No "check for security issues." Just the question a tired senior dev asks before hitting merge at 5 PM on Friday.
Claude Sonnet scored 9/10. GPT-5.5 scored 6/10. But here's what matters: all four models, including the highest-performing one, failed the same task. A one-line change that swapped secrets.compare_digest() for a regular string comparison. The change "looked fine." It "simplified" the code. And every model blessed it.
Why This Matters for Production Work
I think the real insight here isn't that LLMs are bad at security review. It's that LLMs are excellent at pattern matching on keywords but struggle with logic that requires tracing execution through edge cases.
SQL injection? Every model catches it because "SQL" + "string concatenation" = red flag. That's a trained pattern. But a malformed header crashing because a function expected bytes, not Unicode? That requires understanding what happens when your assumptions break. It requires reading between the lines.
This is genuinely dangerous in production. Because the bugs LLMs miss aren't the ones developers write intentionally, they're the ones that hide in legitimate-looking refactors. A new developer cleans up some old code. The LLM says yes. Nobody question it further. Six months later, someone figures out how to exploit it.
My Take: LLMs Are Great Guards, Not Gatekeepers
I'm not anti-LLM for code review. I use them every day. But I've changed how I think about their role. They're excellent at catching:
- Logic that contradicts itself
- Missing null checks in obvious places
- Obvious injection vectors
- Dead code and unused variables
They're genuinely terrible at catching:
- Changes that remove guards while looking like cleanups
- Edge cases that nobody wrote tests for
- Timing or behavioral changes that only surface under load
- Domain-specific knowledge (your app's specific threat model)
The author of the benchmark suggests a follow-up: what if you tell the model exactly what to look for? Instead of "is this safe to merge," ask "does this handle all input types correctly?" I'm curious about that too. But here's what I'd actually want: give the model the test suite. Show it the tests that currently exist for that function. If the refactor causes any tests to fail or changes the behavior in a covered case, that's a strong signal.
Here's What I Changed in Our Workflow
We now use LLM review as a fast first pass, not a gatekeeper. The bot flags obvious issues and style problems. But every security-adjacent change gets manual review, and we're ruthless about it. A "simplification" that touches auth, rate limiting, or comparison logic? That's human-only, no exceptions.
We also started asking LLMs targeted questions instead of open-ended ones: "Does this change the behavior of the permission check?" "Can this function still handle invalid input?" It's not foolproof, but it catches more than blind approval ever did.
What's Your Experience Been?
Have you caught something in code review that an LLM missed? More importantly, have you caught something that an LLM approved that shouldn't have shipped? I'm genuinely interested in whether this pattern holds beyond security issues. Do LLMs miss other kinds of subtle behavioral changes?
I'm not ready to trust LLM code review as a security boundary in production. But I'm also not dismissing it. The tool works best when you know exactly what it's good at and build safeguards around where it fails.
Source: This post was inspired by "The Plausible PR: I Gave 4 LLMs 10 Sneaky Refactors, and They All Missed the Same Bug" by Dev.to. Read the original article