Trust them as a record of what a tool did on someone else's test, not as a ranking. Every vendor-run comparison against other tools in the table below puts its author first, and the same rival gets very different numbers on different vendors' pages, because each page counts something different. Read the method section before the chart, then run your shortlist on your own pull requests.
All pages quoted here were read on 5 October 2026, and the ReviewBench pages on 6 October.
Each page uses its own metric and dataset, so this table holds no scores. The last column is only the name at the top of that page's own table.
| Benchmark | What it counts | Cases | Who judged a match | How rivals were set up | First in its own table |
|---|---|---|---|---|---|
| Greptile , July 2025 | catch rate: a line-level comment naming the faulty code and its impact | 50 bugs, 5 repositories | results "verified against the known bug" | hosted plans, default settings | Greptile |
| Augment , Dec 2025, updated Jun 2026 | precision, recall, F-score against golden comments | 50 PRs, 5 projects, key corrected by Augment | not found in the post | not found in the post | Augment, by F-score |
| Qodo , Feb 2026 | precision, recall, F1 against issues Qodo injected | 100 PRs, 580 issues | an LLM judge | default settings | Qodo |
| Cursor , Apr 2026 | resolution rate: comments addressed by merge | public repositories, 11,419 to 50,310 PRs per tool | an LLM judge | as found in public repositories | Cursor Bugbot |
| Macroscope , Sep 2025 tool benchmark | detection rate of known runtime bugs | 118 bugs, 45 repositories | an LLM, every match checked by hand | default settings, minimum plans | Macroscope |
| Bito , Mar 2026 | coverage of a "Truth Set" | 65 issues, 5 languages | not found in the post | not found in the post | Bito |
| CodeRabbit , Sep 2026 | models inside its own pipeline, not rival tools | 13 cases, plus 44 PRs for speed | "an independent judge", three votes | no rivals | not a tool ranking |
| GitHub ReviewBench , Oct 2026 | grounded precision, recall and F1 against a golden set; findings outside the set are left out | 219 PRs, 187 repositories, 19 languages | an LLM judge, Claude Sonnet 5 | initial entries run by the ReviewBench team on each vendor's public product, June to October 2026; vendors "did not run or verify" them | Copilot Code Review, by grounded F1 on 6 October; changes with the leaderboard |
| Martian Code Review Bench | offline: golden comments; online: what developers fixed | offline 50 PRs; online fresh public PRs | LLM judges | offline: Martian runs the pipeline itself for every tool it publishes; online: as found in public repositories | changes with the leaderboard |
CodeRabbit is easy to follow, because five other vendors include it. Augment's table gives it 36% precision against Augment's golden comments. Greptile gives it a 44% catch rate on Greptile's 50 bugs. Macroscope gives it a 46% detection rate on Macroscope's 118 runtime bugs. Cursor gives it a 48.96% resolution rate on comments in public repositories. Bito writes that "Coderabbit achieved an average coverage of 65.8 percent, coming closest to Bito." Five pages, five things counted, five datasets. Each number can be right on its own page, and none of them belongs in a column next to another.
Greptile built its set from 50 bug-fix pull requests in Sentry, Cal.com, Grafana, Keycloak and Discourse. Augment benchmarks on 50 pull requests from the same five projects and says what it changed: "We expanded and corrected the golden comments by reviewing each PR manually, verifying issues, and validating them against tool outputs." Martian calls its offline set an "Open replication of the code review benchmark used by companies like Augment and Greptile". Under Greptile's scoring, Greptile comes first; under Augment's corrected key, Augment does by F-score. Neither has to be wrong for both to be true. The answer key and the scoring rule are part of the result.
Recall is the share of known issues a tool found. Precision is the share of its comments that were right, and the pages differ on what "right" means.
Against an answer key, a comment is right if it matches an issue someone listed in advance. Qodo defines precision as "the rate of tool-generated comments that correctly correspond to a ground truth issue". A real bug missing from the key counts against the tool.
Against developer behaviour, a comment is right if the author changed the code. Cursor spells out what follows: "52% of the bugs it identified were resolved by the time the relevant PR was merged, indicating the rest were false positives." Under this rule a correct warning the author ignored counts as noise, and a nit the author happened to fix counts as signal. In a catch rate, noise is free: Greptile notes that "false positives, style suggestions, and unrelated comments did not affect the catch rate," and points readers to its case tables to judge noise.
Vendors even disagree on which number matters. Baz: "Precision is the metric we optimize for." Qodo: "precision is a dimension that can be tuned post-processing according to user preference". If your team scrolls past the bot, look at precision against what developers fix; if bugs reach main, look at recall against a key you trust.
The most useful sentences on these pages are often the caveats, and I would read them before any chart. CodeRabbit: "Thirteen cases is a small set, and we say so throughout." On those cases Sonnet 5.5 caught 6 and Sonnet 5 caught 4, a difference of two.
Macroscope kept "only the self-contained runtime bugs, since Macroscope Code Review is designed specifically to detect this type of issue", used the results to fix its own pipeline, while stating that "we have not encoded any of the bugs in our dataset into our pipelines, nor trained any models on this data", and adds: "We acknowledge that the other tools we evaluated did not have the same opportunity to fix any issues that they would have encountered with this exact dataset."
Baz calls many earlier benchmarks in this category "too small to generalize, too static to stay clean, and too tied to a single vendor's scoring choices." CodeAnt is blunter: "Most published comparisons were produced by the vendors themselves. Predictably, every benchmark declared its own tool the winner."
Martian publishes Code Review Bench as open source and, on its site, describes it as "An Unbiased OSS Benchmark For Code Review Agents." Its README names the problem in one line: "Without shared evals for these tools, every company grades its own homework." Its README says Martian runs the pipeline itself for every tool it publishes, a rule added on 25 August 2026, and scores each one offline against 173 golden comments on 50 pull requests (137 in its March data) and online against what developers fixed on fresh ones. That answers question one and, for the pipeline, question five; which configuration of each tool was entered still has to be checked, and the rest still apply. Offline and online rankings can differ because they count different things, which is what puzzled the person who asked this question on r/codereview, and the README adds that "Different LLM judges can score differently," while reporting that "the top 5 tools are identical across all three judges, with most tools varying by at most 2 rank positions." In the dashboard data on 5 October the same five tools led under every judge in the default view, but not in the same order: the judges named two different tools first. Vendors quote it too, each from its own snapshot: in March 2026 Baz reported "Baz ranks #1 in precision in the current results", and CodeAnt that it "ranked #3 globally" by F1. The F1 score CodeAnt reports matches Martian's offline data of 16 March 2026; on 19 March Martian re-scored those results with duplicate findings merged (commit 720e1d3), and first place changed under all three judges.
GitHub's ReviewBench, announced on 5 October, is open source too, with its pull requests, golden findings and judge prompt public. But GitHub runs it, its own Copilot Code Review leads the table, and the announcement says that using it on Copilot code review "has helped us improve the product". Its extraction notes warn that LLM producers "risk biasing the golden set toward the models used to generate them". In the golden files on 6 October, 1,185 of the 2,623 findings marked true positive have a producer field starting with "ccr", the abbreviation the announcement uses for Copilot code review; where producers found the same issue, one copy was kept.
Several of the authors invite it. Bito: "Treat any benchmark as a snapshot, and re-run tests when you see big product updates or if you are about to sign a longer-term contract." CodeRabbit, testing models rather than tools: "Run it on pull requests from your own repositories with known outcomes".
The same steps are how to judge a tool that is in none of these benchmarks, such as Sigma (Mnemoverse, the author's company), a GitHub pull request reviewer in closed beta; it runs on Claude Code and OpenAI models, and its team reviews its own repositories with it every day.
Trust the method sections more than the leaderboards. A vendor benchmark is a fair record of what that vendor chose to measure, on cases it chose, scored by rules it wrote, and the better ones say so themselves. Use the published numbers to pick two or three tools worth a trial, and let your own pull requests decide.