I made CodeRabbit's reviews a third less noisy with an open-source Claude Code skill A developer built pr-proof, an open-source set of three Claude Code skills that validates AI code-review comments against the actual source before acting on them. Run against CodeRabbit's reviews of 50 real pull requests from the Code Review Bench, the tool cut review noise by roughly a third, filtering 300 raised issues down to the 77 real bugs. The author also benchmarked the companion pr-review skill against plain Claude Code, finding them statistically level at 29.8% versus 29.1% F1. AI code review has a noise problem. On a public benchmark of 50 real pull requests, CodeRabbit raised 300 issues. 77 of them were real bugs on the benchmark's list. Most of the rest were noise: things that look like problems but aren't once you read the code. Developers learn to skim review bots, and once you skim, the real bugs slip past you too. I built pr-proof https://github.com/TanayK07/pr-proof , three Claude Code skills that treat every review comment as a claim that has to be proven against the code before anyone acts on it. Run over CodeRabbit's reviews of those 50 PRs, it: // Detect dark theme var iframe = document.getElementById 'tweet-2106061332357529939-740' ; if document.body.className.includes 'dark-theme' { iframe.src = "https://platform.twitter.com/embed/Tweet.html?id=2106061332357529939&theme=dark" } Here are real comments CodeRabbit left on a Grafana PR, and what checking them against the code showed: " promResponse.data.groups.at 0 ?.rules.map will throw if groups is empty." It won't. Optional chaining short-circuits the whole rest of the chain: if at 0 returns undefined , .rules.map ... is never evaluated. " useCanSilence invokes hooks before its early return when rule is undefined." That order is required. The Rules of Hooks say hooks must run unconditionally, before any early return. " handleDelete expects RulerRuleDTO but call sites pass EditableRuleIdentifier ." Both call sites pass zero-argument closures RuleActionsButtons.V2.tsx:99 . Each one sounds plausible and falls apart once you read the code, and reading the code is exactly the step review bots skip and leave to you. The core skill, pr-comment-validation , takes review comments from anyone a human, CodeRabbit, Copilot, another agent and investigates each one: Nothing is changed or posted until every comment has a verdict. Two more skills build on it: pr-validation does the whole loop. It checks out the PR in a worktree, validates every comment in parallel, shows you the verdicts, applies the fixes you approve, and replies on each thread. pr-review writes its own review, then has independent subagents try to disprove each finding before anything is posted. Install it in Claude Code: /plugin marketplace add TanayK07/pr-proof /plugin install pr-proof@pr-proof Then, on any PR with review comments: "are the review comments on PR 123 valid?" I used Code Review Bench https://github.com/withmartian/code-review-benchmark , an open benchmark with 50 real PRs from Sentry, Grafana, Keycloak, Discourse and Cal.com. Each PR comes with a human-written list of the real issues. It already holds every major tool's reviews, scored by Claude Opus 4.5 as the judge. That made the filter test clean: Each run was a headless Claude Code session with no network, no gh and no curl , so it couldn't read the original PR discussions the answers came from. I also benchmarked pr-review as a reviewer in its own right, against plain Claude Code on the same model. It came out statistically level: 29.8% vs 29.1% F1. Two things I learned along the way: pr-comment-validation uses fixed the harm, but only added about 1.6 points. I also caught a bug in my own scoring halfway through. On 37 PRs the benchmark had silently fallen back to judging whole comments instead of individual issues, which made pr-review look like it scored 45%. It didn't. The harness now refuses to print results if that happens. Everything is in the repo: the skills, the benchmark harness, every review and verdict per PR, and the scripts that turn them into these numbers. bench/README.md If you run it on your own PRs, I'd love to hear where it gets things wrong.