AI Code Review Is Not a Replace-and-Forget Tool, And This Benchmark Proves It
I got burned by an LLM code review last month. We integrated Claude into our PR workflow thinking it would catch the obvious stuff, unused imports, basic logic errors, that kind of thing. A junior dev submitted a change that looked clean on the surface: consolidating an authentica...