{"slug": "reviewbench-an-open-benchmark-for-ai-code-review", "title": "ReviewBench: An open benchmark for AI code review", "summary": "GitHub launched ReviewBench, an open benchmark for AI code review agents built on 219 pull requests from 187 public open source repositories across 19 languages, modeled after an analysis of 103.9 million GitHub pull requests. The benchmark uses a multi-source golden set, a consistent evaluation rubric independently validated by senior engineers, and six metrics in two families, and GitHub says its offline evaluation of Copilot code review has become more effective at anticipating the direction of production experiments. The dataset is publicly available and teams can onboard their own code review systems and submit results.", "body_md": "### \n[Michelle Zhou](https://github.blog/author/miichellezhou/)\n\nMichelle develops and evaluates agentic AI systems for code review, focusing on repository-level context retrieval, review and fix quality, and rigorous benchmarking of AI code reviewers.\n\nWe’re launching ReviewBench, a benchmark for code review agents built on representative GitHub pull requests, multi-source ground truth, calibrated evaluation, and production-aligned metrics.\n\nAgentic code review is becoming an essential piece of how development happens. It helps you inspect pull requests, catch issues, and decide what deserves attention before code ships.\n\nBut the quality of existing AI reviewers can be hard to measure, and you need to know the strengths of a reviewer before you know if it will help you. Some reviewers surface more issues, some produce less noise, and some are stronger at catching critical problems while others surface smaller improvements, too. You may need code review to do different things within your workflow.\n\nThat makes it important to understand how reviewers actually compare: what different systems catch, what they miss, and the tradeoffs they make. A good code review benchmark should reflect the diversity of real pull requests, capture a broad set of review findings, and support meaningful breakdowns by severity, category, and precision-recall preferences. For teams building code review agents, the benchmark should also provide an offline signal that reliably tracks whether changes are likely to improve the experience in production. Existing benchmarks often make tradeoffs between label quality, coverage, and how well they represent real-world code review, leaving a gap for a rigorous and reproducible evaluation methodology that brings these pieces together.\n\nWe built **ReviewBench****, a new code review offline benchmark**, to address that gap, and it is available for you to use today. It follows the language, repo size, and size distribution of pull requests, modeled after over 100 million real pull requests on GitHub. It uses a multi-source golden set and a consistent evaluation rubric and has been independently validated by senior engineers. Just as important, with the help of ReviewBench, our offline evaluation of Copilot code review (CCR) has become more effective at anticipating the direction of production experiments, giving us greater confidence that measured improvements reflect meaningful gains for users.\n\nIn this post, we’ll walk through how ReviewBench is constructed, how it establishes reliable ground truth and scoring, and how to onboard your own code review system and submit results.\n\nOur benchmark is built around five principles:\n\n**1. Representative pull requests, not a demo set**\n\nWe analyzed 103.9 million GitHub pull requests to characterize the real-world distribution of code review workloads. ReviewBench contains 219 pull requests from 187 public open source licensed repositories spanning 19 languages, with its language and repository-size distributions closely matching GitHub overall. The complete benchmark dataset is publicly available.\n\nWe make one deliberate adjustment to this distribution: while language and repository size mirror GitHub directly, pull request size is weighted toward the reviewable middle and tail. This reduces the overrepresentation of tiny, single-file changes while preserving more substantive, multi-file pull requests where review quality matters most.\n\nQuick corpus snapshot:\n\n**2. Broad ground truth discovery, independently judged**\n\nNo single reviewer, whether human or model, can identify everything worth finding in a pull request. To build a broader and more reliable golden set for ground truth findings, we follow a three-stage process:\n\n**3. Metrics that measure both known and newly discovered issues**\n\nMost benchmarks report precision and recall against a fixed golden set. ReviewBench reports six metrics in two families:\n\nThat distinction becomes more important as review agents become more capable. A fixed golden set inevitably becomes incomplete as systems discover issues its creators did not anticipate. Augmented metrics let ReviewBench recognize that behavior rather than automatically penalizing it. Because augmented recall expands the denominator based on what each agent discovers, we use grounded recall as the headline cross-system comparison and augmented metrics as an additional per-system diagnostic.\n\n**4. Configurable evaluation for different review preferences**\n\nThere is no single universally optimal review experience. Some developers may want to focus only on critical issues, while others also value lower-severity, non-breaking findings. Some prefer broader coverage, while others prioritize precision and minimal noise. Others may have specialized needs, such as security- or privacy-focused review.\n\nReviewBench lets results be sliced by severity and category, while precision and recall capture different operating preferences. Users can also adjust β in the Fβ score to place more weight on recall for broader coverage or precision for lower noise. As these preferences change, the leaderboard is re-ranked accordingly, helping users identify the systems that best match their review priorities.\n\n**5. Internally audited and reproducibly evaluated**\n\nBefore release, we asked senior engineers who had not participated in building the benchmark dataset to independently re-label every ground-truth finding from scratch. Their true/false-positive judgments agreed with ReviewBench 96.6% of the time. We version the benchmark dataset, judge, and matcher used in every evaluation, so results can be compared under the same benchmark configuration and revalidated when the benchmark changes. We also publish the validation methodology, agreement measurements, and known threats to validity, so readers can see how benchmark quality is assessed and where uncertainty remains.\n\nReviewBench’s research preview version is now available through [the ReviewBench website](https://review-bench.ai/), where you can explore the full benchmark, compare code review agents, and bring your own agent to evaluate and iterate.\n\nWith ReviewBench, you can:\n\nWe have used ReviewBench to evaluate [Copilot code review (CCR)](https://docs.github.com/en/copilot/how-tos/use-copilot-agents/request-a-code-review/use-code-review) across successive iterations, giving us a consistent way to measure progress, catch regressions, and prioritize promising changes. Over time, this has helped us improve the product. One of the most valuable benefits of ReviewBench is that it provides an early offline signal of how a change to the product is likely to perform in production. Across experiments evaluated with ReviewBench before A/B testing, offline changes have consistently pointed in the same direction as what we see later in production.\n\nA recent lite-tier experiment provides a concrete example of this broader pattern. We introduced a multi-model ensemble review that combines several independent model runs into a single review rather than relying on a single run. ReviewBench predicted higher precision, recall, and comment volume, along with lower cost per review.\n\nTo compare offline and production results, we use corresponding online signals. Addressed rate, our online counterpart to precision, is the percentage of CCR comments that an LLM determines prompted a developer to make a corresponding code change, based on the diff, thread, reactions, resolution state, and post-review code. For recall, we measure how much additional human review is still needed.\n\nThe online A/B test moved in the same direction as ReviewBench predicted: addressed rate (precision) rose 8.0%, recall rose 13.6%, and comment volume rose 61%, while cost per review fell 8.0%, all relative to the production control.\n\nComment volume alone, however, does not capture comment quality. More critical findings mean something very different from low-severity nits. ReviewBench’s severity-level evaluation captured this too: it predicted a 227% increase in critical comments, compared with 262% online, along with the same broader shift toward more moderate comments and fewer nits.\n\nThis gives us a fast and repeatable signal before running production experiments. Online experiments remain the ultimate measure of user impact, but ReviewBench gives us greater confidence in which changes are worth taking there.\n\nWe invite you to explore [ReviewBench](https://review-bench.ai/), evaluate your own system, challenge our assumptions, and help us improve the benchmark. We’re excited to collaborate with researchers and practitioners to make code review evaluation more open, reliable, and useful—and ultimately help move AI code review forward.\n\nReviewBench was a team effort across GitHub and Microsoft. We’re grateful to the researchers and engineers who built it: those who designed the methodology, curated the pull requests, built the golden set and the evaluation pipeline, and made the benchmark something anyone can run.\n\nLearn to direct AI agents, critically review their output, and keep technical judgment at the center of your workflow.\n\nDescribe the interface you need in plain English, then let the agent build a live surface you can both use and update—so you spend less time adapting to tools and more time getting work done.\n\nWhat is a developer to do when they need something more tangible than a chat box? Enter canvases.", "url": "https://wpnews.pro/news/reviewbench-an-open-benchmark-for-ai-code-review", "canonical_source": "https://github.blog/ai-and-ml/github-copilot/reviewbench-an-open-benchmark-for-ai-code-review/", "published_at": "2026-10-05 15:59:40+00:00", "updated_at": "2026-10-05 16:16:49.866297+00:00", "lang": "en", "topics": ["ai-tools", "developer-tools", "artificial-intelligence", "ai-products", "ai-research"], "entities": ["GitHub", "ReviewBench", "Copilot code review", "Michelle Zhou"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/reviewbench-an-open-benchmark-for-ai-code-review", "markdown": "https://wpnews.pro/news/reviewbench-an-open-benchmark-for-ai-code-review.md", "text": "https://wpnews.pro/news/reviewbench-an-open-benchmark-for-ai-code-review.txt", "jsonld": "https://wpnews.pro/news/reviewbench-an-open-benchmark-for-ai-code-review.jsonld"}}