# I made CodeRabbit's reviews a third less noisy with an open-source Claude Code skill

> Source: <https://dev.to/tanayk07/i-made-coderabbits-reviews-a-third-less-noisy-with-an-open-source-claude-code-skill-eho>
> Published: 2026-10-03 05:23:38+00:00

AI code review has a noise problem. On a public benchmark of 50 real pull requests, CodeRabbit raised 300 issues. 77 of them were real bugs on the benchmark's list. Most of the rest were noise: things that look like problems but aren't once you read the code.

Developers learn to skim review bots, and once you skim, the real bugs slip past you too.

I built [pr-proof](https://github.com/TanayK07/pr-proof), three Claude Code skills that treat every review comment as a claim that has to be proven against the code before anyone acts on it. Run over CodeRabbit's reviews of those 50 PRs, it:

// Detect dark theme var iframe = document.getElementById('tweet-2106061332357529939-740'); if (document.body.className.includes('dark-theme')) { iframe.src = "https://platform.twitter.com/embed/Tweet.html?id=2106061332357529939&theme=dark" }

Here are real comments CodeRabbit left on a Grafana PR, and what checking them against the code showed:

*"`promResponse.data.groups.at(0)?.rules.map()` will throw if groups is empty."*

It won't. Optional chaining short-circuits the whole rest of the chain: if `at(0)` returns `undefined`, `.rules.map(...)` is never evaluated.

*"`useCanSilence` invokes hooks before its early return when rule is undefined."*

That order is required. The Rules of Hooks say hooks must run unconditionally, before any early return.

*"`handleDelete` expects `RulerRuleDTO` but call sites pass `EditableRuleIdentifier`."*

Both call sites pass zero-argument closures (`RuleActionsButtons.V2.tsx:99`).

Each one sounds plausible and falls apart once you read the code, and reading the code is exactly the step review bots skip and leave to you.

The core skill, `pr-comment-validation`, takes review comments from anyone (a human, CodeRabbit, Copilot, another agent) and investigates each one:

Nothing is changed or posted until every comment has a verdict.

Two more skills build on it:

`pr-validation` does the whole loop. It checks out the PR in a worktree, validates every comment in parallel, shows you the verdicts, applies the fixes you approve, and replies on each thread.`pr-review` writes its own review, then has independent subagents try to disprove each finding before anything is posted.
Install it in Claude Code:

```
/plugin marketplace add TanayK07/pr-proof
/plugin install pr-proof@pr-proof
```

Then, on any PR with review comments: *"are the review comments on PR #123 valid?"*

I used [Code Review Bench](https://github.com/withmartian/code-review-benchmark), an open benchmark with 50 real PRs from Sentry, Grafana, Keycloak, Discourse and Cal.com. Each PR comes with a human-written list of the real issues. It already holds every major tool's reviews, scored by Claude Opus 4.5 as the judge.

That made the filter test clean:

Each run was a headless Claude Code session with no network, no `gh` and no `curl`, so it couldn't read the original PR discussions the answers came from.

I also benchmarked `pr-review` as a reviewer in its own right, against plain Claude Code on the same model. It came out statistically level: 29.8% vs 29.1% F1.

Two things I learned along the way:

`pr-comment-validation` uses fixed the harm, but only added about 1.6 points.
I also caught a bug in my own scoring halfway through. On 37 PRs the benchmark had silently fallen back to judging whole comments instead of individual issues, which made pr-review look like it scored 45%. It didn't. The harness now refuses to print results if that happens.

Everything is in the repo: the skills, the benchmark harness, every review and verdict per PR, and the scripts that turn them into these numbers.

`bench/README.md`
If you run it on your own PRs, I'd love to hear where it gets things wrong.
