Devlery
Blog/Microsoft

GitHub Opened a Benchmark for AI Code Review, and the Leader Finds 26% of Known Defects

GitHub published ReviewBench on October 5, scoring code review agents against 219 public PRs. Copilot Code Review leads at 87.8% precision but 26.0% recall, and on critical defects Codex scores 81 to Copilot 63.

GitHub Opened a Benchmark for AI Code Review, and the Leader Finds 26% of Known Defects
AI 요약
  • GitHub opened ReviewBench, which scores review agents against 219 public PRs.
  • Copilot leads: 87.8% of its comments are real defects, but it finds only 26.0% of known ones.
  • On critical defects Codex scores 81 and Copilot 63.

GitHub published ReviewBench on October 5, 2026. It is an open test of how many real defects an AI code review bot catches on actual pull requests, and it reports two numbers: how many of the problems a human reviewer would have flagged the bot found, and how many of its comments were wrong.

The release is unusually complete. 219 PRs drawn from 187 public repositories, a golden answer set for each PR, the grading rubric, the judge model configuration, and a script for running your own agent against the corpus are all in the repository under the MIT license. Nineteen languages are represented, TypeScript most heavily at 68 PRs. A leaderboard went live at review-bench.ai alongside it.

Precision is flat across the field. Recall is where they split

AI code review tools have multiplied since the market took shape around Qodo's $70M round, but almost no public numbers measured them against the same ruler. Here are the seven default rows on the leaderboard, as of October 9, 2026.

RankReviewerPrecisionRecallRun date
1Copilot Code Review (Balanced)87.8%26.0%Oct 1
2Devin AI84.0%23.8%Sep 28
3Qodo85.3%22.1%Sep 28
4Codex (GPT-5.6 Sol, Max)86.0%18.5%Oct 8
5Cubic85.5%16.3%Jun 29
6Greptile86.1%16.2%Jun 16
7Cursor87.7%9.6%Sep 27

Both numbers need unpacking. Precision is the share of the bot's comments that pointed at a real problem; higher means less noise. Recall is the share of the already-known real problems in that PR the bot surfaced; higher means fewer misses.

Precision clusters between 84.0% and 87.8% across all seven tools, a spread of under four points. Recall runs from 26.0% at the top to 9.6% at the bottom, close to a threefold gap. Across all 28 leaderboard entries the floor is 5.7%, from Codex driven by GPT-5.6 Terra at Low effort.

Today's AI code reviewers are tuned to stay quiet. When they speak they are usually right, but three of every four known problems pass without comment. The leader is no exception, and that one line is the first thing to take from the board.

The leaderboard also carries Augmented precision and Augmented recall columns, and those numbers look better. Copilot's Augmented recall is 34.7%, because a comment absent from the golden set still earns credit if the judge agrees it is valid. The methodology document says explicitly not to rank tools by that column. The more novel findings an agent emits, the more the denominator grows with it, so an equally capable agent scores higher simply by talking more. The column designated for cross-tool comparison is the recall in the table above.

The overall leader is not the leader on critical defects

Split the same leaderboard by defect severity and the order inverts. The table below is F1 by severity, the single number that folds precision and recall together.

ReviewerCriticalMediumLow
Codex (GPT-5.6 Sol, Max)816125
Copilot Code Review (Balanced)636038
Devin AI504017
Qodo454826
Cubic414337
Greptile363730

Codex sits fourth overall and first on critical defects at 81, against 63 for Copilot Code Review. On low-severity defects the order flips: Copilot 38, Codex 25. Copilot leads the overall table not because it catches more serious bugs but because it picks up the minor ones too.

The category breakdown splits the same way. On recall measured over critical and medium defects only, Codex leads on correctness at 50 and reliability at 49, against Copilot's 43 and 37. Copilot leads by a wide margin on missing tests at 54 and maintainability at 23, against Codex's 26 and 8. Want a reviewer that notices the test you forgot to write, pick Copilot. Want one that notices the logic is wrong, pick Codex.

Reading only the overall rank erases that distinction entirely. The question to pick on is not the ranking but which defect hurts most when your repository ships it.

How far to trust these numbers

ReviewBench recall answers "how many of the known defects did it find," so what sits in the golden set determines what the number means.

I tallied the 219 golden-set files in the repository directly. Of 4,632 total candidate findings, 2,623 are labelled true defects, and they come from here.

LLM reviewers (Sonnet 4.6 and others)
1,391 · 53.0%
Copilot code review
1,185 · 45.2%
Human review comments
39 · 1.5%
Static analysis (Semgrep)
8 · 0.3%

Counted from the producer field of entries where tp_fp is tp, across the 219 golden/*.json files in review-bench/ReviewBench

Ninety-eight percent of the golden set was written by LLM reviewers. Only 39 entries came from human review comments. Of the 191 comments humans left on these PRs, 152 were classified as out-of-scope observations or nitpicks and dropped. The announcement's phrase "human-audited golden set" does not mean humans wrote the findings; it means humans audited the labels a classifier assigned. In that audit humans and the classifier agreed on the true-positive call 96.6% of the time, and humans corrected 47 entries the classifier had mislabelled.

The set is also not a collection of hard problems. Of the 2,623 true defects, 1,311 are rated easy and 1,073 medium, leaving 239 hard, or 9.1%. Severity skews the same direction, with 1,551 rated low, more than half. The 26.0% at the top of the board was measured against that set.

One more thing. Of the 2,623 true defects, 1,185 (45.2%) were produced by producers tagged ccr:. GitHub abbreviates Copilot code review as CCR in the announcement. The product sitting first on the leaderboard generated 45% of the answer key.

GitHub does not hide this. Chapter 8 of the methodology document names producer bias and classifier dependence outright:

The classifier prompt encodes our working group's taste, and a different team configuring a classifier under the same methodology could produce different scores for the same agent.

The leaderboard's own notice says the same. The initial entries were produced by the ReviewBench team running each vendor's product themselves, vendors did not validate them, and the results "may not predict performance on your code." That is also why the run dates in the table matter. Cubic was measured on its June 29 product and Greptile on June 16, while the first-place Copilot entry is from October 1.

Set this against the other public yardstick and the order changes. Martian's Code Review Bench tracks whether a bot's suggestion on a real public PR actually led to a code change. On its online leaderboard, which has graded 13,274 PRs in the past month, GitHub Copilot ranks fifth.

Top five of the Martian Code Review Bench online leaderboard, with Cubic Dev AI first and GitHub Copilot fifth

Copilot is fifth on the offline leaderboard too, scoring 58.0% F2 on a separate set of 50 hard-to-find bugs, behind Qodo Deep at 65.1%. Cubic and Greptile, fifth and sixth on ReviewBench, sit first and second on Martian's online board. The two benchmarks measure different things by construction: ReviewBench compares against a golden set on a fixed corpus of 219 PRs, while Martian watches tens of thousands of live PRs to see whether developers accepted the suggestion. A vendor's own measurement diverging from a neutral one is the same pattern model benchmarks have already been through.

Running it yourself

There are two paths. Cloning the repository and running locally requires no approval; getting onto the leaderboard goes through a submission process.

ItemDetail
Who can use itAnyone. Dataset, golden set, grading rubric, judge prompts, and run scripts are MIT licensed
PriceDataset free. Tuning and test runs are on your own model keys and compute. ReviewBench covers only the official judge cost for a final leaderboard submission
Regional availabilityNo geographic restriction stated. The repository clones anywhere, and leaderboard submission needs a GitHub account
Local requirementsDocker, git, jq
Leaderboard requirementsPush your agent container to GHCR pinned by digest, register at review-bench.ai/submit with a GitHub account, then wait for a maintainer to merge
Corpus contributionsNot accepted. The corpus is closed to new PRs

A local run is three lines from the README.

git clone https://github.com/review-bench/ReviewBench && cd ReviewBench
scripts/try-agent.sh my-reviewer:dev -e MY_API_KEY             # 25-PR test set
scripts/try-agent.sh my-reviewer:dev --set full -e MY_API_KEY  # all 219 PRs

An adapter has to read one PR and write one findings file. The README notes that one open source reviewer connected in roughly 90 lines. Be aware that try-agent.sh only confirms your agent runs to completion and emits well-formed output; it does not score anything. Scoring is a separate npm run judge pass on your own model keys.

If your team already has a review bot wired into internal PRs, run it against the 25-PR test set first and look at the severity breakdown before anything else. A low score on critical defects means you should not be cutting human review on the strength of that bot, whatever its overall rank says.