# GitHub ReviewBench: Copilot Is #1 on Its Own Test, #5 Elsewhere

> Source: <https://byteiota.com/github-reviewbench-copilot-ranking/>
> Published: 2026-10-08 21:08:47+00:00

GitHub dropped [ReviewBench](https://github.blog/ai-and-ml/github-copilot/reviewbench-an-open-benchmark-for-ai-code-review/) last week: an open benchmark for AI code review agents, built on 219 real pull requests across 19 programming languages. GitHub Copilot landed at number one. Also grading the exam: GitHub.

An [independent benchmark run by Martian](https://codereview.withmartian.com/) tells a different story. On Martian’s offline leaderboard, Copilot sits fifth — behind Qodo Deep, Cubic, Augment, and one more. On the online leaderboard, it climbs to fourth. Either way, the tool that ranked first on GitHub’s own test doesn’t crack the top three anywhere else.

## What ReviewBench Actually Measures

ReviewBench is technically sound on paper. GitHub analyzed 103.9 million pull requests to build a representative corpus, drew ground truth from multiple sources — human reviewers, static analysis tools, and frontier LLMs — and had senior engineers independently re-label every finding, achieving 96.6% agreement. That’s the kind of rigor you want in a benchmark.

Where it gets complicated: roughly 45% of the ground truth findings originated from Copilot itself. The benchmark’s answer key is built, in part, from the output of the product being evaluated. GitHub didn’t hide this — it’s documented in the methodology. But it means the test has Copilot’s fingerprints on it before anyone submits an answer.

## The Rankings, Side by Side

| Benchmark | Operator | #1 Tool | Copilot Rank | 
|---|---|---|---|
| GitHub ReviewBench | GitHub (vendor) | Copilot — 40.1% F1 | 1st | 
| Martian Bench (online) | Martian (independent) | Cubic — 64.9% F1 | 4th (60.9% F1) | 
| Martian Bench (offline) | Martian (independent) | Qodo Deep | 5th (58% F2) | 

The raw scores aren’t directly comparable — different corpora, different graders, different severity weighting. But the ranks speak clearly enough.

## Why the Gap Exists

GitHub’s benchmark is offline: present the AI with a PR, check its findings against a fixed golden set, score once. Martian adds an online dimension: track whether real developers actually implement the suggestions in real pull requests. That second measurement is unforgiving. A finding that’s technically correct but contextually useless doesn’t get credit because no developer acts on it.

GitHub also ran every initial leaderboard entry itself — Copilot was tested in October, competitors in June. The methodology is open-source, but no vendor has independently verified their own results on the platform yet. For a benchmark that bills itself as open, the initial dataset is remarkably one-sided.

## This Is What Every Vendor Does

The frustration in the developer community isn’t specific to GitHub. One analysis found the same tool — CodeRabbit — scored anywhere from 36% to 65.8% across five vendor-published benchmarks. As one developer put it bluntly: “Every vendor-created benchmark places its own tool first. Predictably.”

[The New Stack noted](https://thenewstack.io/github-reviewbench-code-review/) that GitHub both develops ReviewBench and evaluates Copilot within it — a structural conflict that doesn’t require bad faith to produce distorted results. The incentives do the work. When your revenue depends on your product looking good, your benchmark will find a way to make it look good, even with the best intentions.

Martian’s approach of measuring developer acceptance rates in production sidesteps this entirely. You can’t fake whether a developer merged the fix.

## What to Do Instead

If you’re choosing a code review tool for your team right now, vendor benchmarks are entertainment. The only test that matters is this: pick 20 merged commits from your repo where a bug was later caught in production or via an incident post-mortem. Run each candidate tool against those PRs blind. Check how many real bugs it found. Check how much noise it generated. That’s your benchmark — calibrated to your language, your codebase, your team’s threshold for review fatigue.

ReviewBench is a legitimate contribution to the field. An open corpus, published methodology, independent auditing — that’s better than most. But “better than most vendor benchmarks” is a low bar, and Copilot ranking first on a test GitHub built and populated shouldn’t change your tooling decisions. [The skepticism from the developer community is warranted.](https://dev.to/izgorodin/should-we-trust-ai-code-review-benchmarks-3p6n)
