# How we benchmark AI code reviewers: per-bug scoring on real PRs (https://tenki.cloud/blog/benchmarking-ai-code-reviewers)

> For the complete documentation index, see [llms.txt](https://tenki.cloud/llms.txt)

- Author: Hayssem Elsayed (Product)
- Published: May 19, 2026
- Updated: Sep 22, 2026
- Category: [AI Code Review](https://tenki.cloud/blog/category/ai-code-review)
- Reading time: 7 min

How Tenki's code review benchmark scores seven AI reviewers bug by bug on 50 real pull requests, and what per-bug scoring does and doesn't show.

The useful question about an AI code reviewer is how often it catches the bug: the one a careful human would have spotted in the diff. Did the reviewer find it, point at the right line, and explain why it matters?

We built a benchmark to answer that across seven tools. This post covers how it works, why we score each bug separately instead of each pull request, and where the method has limits. The live leaderboard, with per-repository and per-bug breakdowns, is at [/benchmarks/code-reviewer](https://tenki.cloud/benchmarks/code-reviewer). The numbers below are from the run published on May 20, 2026.

In short: across 50 pull requests with 122 ground-truth bugs, Tenki caught 84 (68.9% recall), 1.9x the next tools, Greptile and Devin, at 44 each. Tenki's precision is 0.30: it catches more bugs by commenting more, which the precision section covers.

## The corpus

The 50 pull requests come from the public benchmark Greptile published in [July 2025](https://www.greptile.com/benchmarks): real, merged bug fixes from five open-source codebases, replayed in reverse so the bug is back in the diff. Ten PRs per repository:

* [cal.com](https://github.com/codereview-benchmark/cal.com/pulls) (TypeScript)
* [sentry](https://github.com/codereview-benchmark/sentry/pulls) (Python)
* [grafana](https://github.com/codereview-benchmark/grafana/pulls) (Go)
* [keycloak](https://github.com/codereview-benchmark/keycloak/pulls) (Java)
* [discourse](https://github.com/codereview-benchmark/discourse/pulls) (Ruby)

Each tool reviewed with its default configuration and full repository context. Nobody added custom rules or per-repo tuning, and that includes Tenki.

The reviews come from two runs:

* For CodeRabbit, Copilot, Cursor, Graphite and Greptile, we re-scored the review comments from Greptile's July 2025 run, which are still public in the [ai-code-review-evaluation](https://github.com/ai-code-review-evaluation) GitHub org.
* For Tenki and Devin, we replayed the same diffs into forks under the [codereview-benchmark](https://github.com/codereview-benchmark) org and collected their reviews in April 2026.

## Ground truth: 122 bugs

Greptile's benchmark scores one bug per PR. We score every bug in the diff, and many of these PRs contain more than one. Each ground-truth finding has a written description: where it is, what triggers it, and why it breaks.

The 122 findings are the regressions each merged fix addressed, plus other bugs in the same diffs that we identified while building the ground truth. Most PRs have one or two findings; a few have seven to ten. Every finding is listed in the case library on the [benchmark page](https://tenki.cloud/benchmarks/code-reviewer), with each tool's verdict.

## How the data was produced

For Tenki and Devin, an automated pipeline worked through the corpus PR by PR, with several agents running in parallel:

1. Replay the bug-introducing diff into a fresh branch on the fork.
2. Open a pull request against the fork's default branch.
3. Wait for each reviewer bot to comment, and give up after a timeout if it stays silent.
4. Collect every review comment on the PR, attributed by bot login.

For the other five tools, the pipeline collected the existing review comments from the July 2025 PRs.

Judging is the same for all seven. For each of the 122 findings, three LLM judges read the tool's comments and vote independently on whether that finding was caught. The published verdicts record an OpenAI Codex model and a Claude model on every finding, with the third seat filled by Gemini on about half the verdicts and Claude Opus on the rest. Two of three votes decide.

A finding counts as caught only when a line-level comment identifies the faulty code and explains the impact. Summary-only mentions, generic "consider adding a test" remarks and comments on the wrong file don't count.

## Why we score per bug instead of per PR

Greptile's original results, like several other published reviewer benchmarks, lead with a per-PR catch rate: did the tool flag at least one real bug in the PR. On that metric, a reviewer that catches the easy bug in a nine-bug PR and misses the other eight gets the same credit as one that finds all nine.

Scored per PR, the top four look close:

| Tool       | PRs caught |
| ---------- | ---------- |
| Tenki      | 32 / 50    |
| Greptile   | 31 / 50    |
| Devin      | 30 / 50    |
| Cursor     | 30 / 50    |
| CodeRabbit | 24 / 50    |
| Copilot    | 24 / 50    |
| Graphite   | 4 / 50     |

Scored per bug, they separate:

| Tool       | Bugs caught (of 122) |
| ---------- | -------------------- |
| Tenki      | 84                   |
| Greptile   | 44                   |
| Devin      | 44                   |
| Cursor     | 39                   |
| CodeRabbit | 35                   |
| Copilot    | 30                   |
| Graphite   | 4                    |

Per bug, Tenki catches 1.9x as many as the next tools. The difference is in the PRs with several bugs: a tool that flags one issue and stops gets full per-PR credit but only part of the per-bug credit.

## Recall, precision and F1

Recall alone can be gamed: a reviewer that comments on every line catches everything. So the benchmark also counts each tool's line comments that don't match any ground-truth finding.

* **Recall** is the share of real bugs the reviewer caught.
* **Precision** is caught bugs divided by caught bugs plus unmatched comments.
* **F1** is the harmonic mean of the two. It doesn't let one number make up for the other linearly: 0.9 precision with 0.3 recall gives an F1 of 0.45, not 0.6.

Devin and Cursor are coding agents: they review their own work in the loop rather than reviewing a PR diff after the fact. As on the benchmark page, they're shown separately for reference, next to the five dedicated PR reviewers.

Dedicated PR reviewers:

| Tool       | Recall | Precision | F1   |
| ---------- | ------ | --------- | ---- |
| Tenki      | 0.69   | 0.30      | 0.42 |
| CodeRabbit | 0.29   | 0.25      | 0.27 |
| Greptile   | 0.36   | 0.16      | 0.22 |
| Copilot    | 0.25   | 0.19      | 0.21 |
| Graphite   | 0.03   | 0.50      | 0.06 |

Coding agents, for reference:

| Tool   | Recall | Precision | F1   |
| ------ | ------ | --------- | ---- |
| Devin  | 0.36   | 0.47      | 0.41 |
| Cursor | 0.32   | 0.51      | 0.39 |

Tenki leads on recall by a wide margin and has the highest F1, with lower precision than Devin and Cursor. Devin and Cursor reach a similar F1 from the other direction, posting fewer comments and catching fewer bugs.

## By severity

Each finding carries its own severity. Across the 122: 7 critical, 53 high, 52 medium, 10 low.

| Tool       | Critical (7) | High (53) | Medium (52) | Low (10) |
| ---------- | ------------ | --------- | ----------- | -------- |
| Tenki      | 4            | 39        | 36          | 5        |
| Greptile   | 5            | 17        | 19          | 3        |
| Devin      | 2            | 24        | 15          | 3        |
| Cursor     | 1            | 18        | 18          | 2        |
| CodeRabbit | 3            | 13        | 17          | 2        |
| Copilot    | 3            | 13        | 12          | 2        |
| Graphite   | 0            | 1         | 2           | 1        |

The critical bucket is too small to rank on: seven bugs, with Greptile at 5 and Tenki at 4. The high bucket has enough samples to mean something. Tenki caught 39 of 53 (74%); Devin was next at 24 (45%).

## Unmatched comments

Tenki posted 289 line comments. 197 of them matched no ground-truth finding; the other 92 matched one of the 84 bugs it caught (some bugs drew more than one comment). Greptile posted a similar 283 comments, with 233 unmatched, and caught 44 bugs.

An unmatched comment is not always a bad one. Some are nitpicks or hallucinated issues. Others point at a real problem the ground truth doesn't list, such as an edge case in a sibling function. We count them all against precision because a "plausible but unverified" bucket would need hand-labeling, and the strict definition is easier to defend.

In practice, Tenki sits at the high-recall end. It catches more bugs and posts more comments to do it. Devin and Cursor sit at the high-precision end with fewer comments and fewer catches.

## What's public

The replayed PRs and every tool's review comments are on GitHub: Tenki's and Devin's in the [codereview-benchmark](https://github.com/codereview-benchmark) org, the other five in [ai-code-review-evaluation](https://github.com/ai-code-review-evaluation). The [benchmark page](https://tenki.cloud/benchmarks/code-reviewer) lists every finding with each tool's verdict, the judge vote count and a link to the review. The judge prompts and the full ground-truth descriptions are not published.

## Using Tenki Code Reviewer

Tenki Code Reviewer is a GitHub App that reviews pull requests for $1.00 per review from workspace credits. If the comment volume above is more than your team wants, the [severity threshold](https://tenki.cloud/docs/reviewer/settings.md) filters which findings appear in the review; all levels are on by default. The [quickstart](https://tenki.cloud/docs/reviewer/quickstart.md) covers setup.