GitHub has launched ReviewBench, an open benchmark for evaluating AI code review agents. It is built on representative GitHub pull requests and includes multi-source ground truth, calibrated evaluation, and production-aligned metrics.
Key facts
| Fact | Detail | The source says |
|---|---|---|
| Purpose | Evaluating AI code review agents | “We’re launching ReviewBench, a benchmark for code review agents built on representative GitHub pull requests, multi-source ground truth,…” |
| Availability | Now available through the ReviewBench website | “ReviewBench’s research preview version is now available through the ReviewBench website” |
| Features | Full benchmark dataset, leaderboard, and self-serve runner | “The complete ReviewBench dataset is publicly available, including the pull requests, findings, labels, severity and category annotations.” |
What happened
GitHub has introduced ReviewBench, an open benchmark designed to evaluate AI code review agents. This benchmark is built on representative GitHub pull requests and includes multi-source ground truth, calibrated evaluation, and production-aligned metrics. ReviewBench allows users to explore the full benchmark dataset, compare systems on a leaderboard, and bring their own agents to evaluate and iterate. The research preview version of ReviewBench is now available through the ReviewBench website.
What to weigh
- Scores remain private until a maintainer reviews and approves the submission.
- Scores are published to the leaderboard only if they outperform the agent's current leaderboard score or if this is the agent's first leaderboard entry.
Source: GitHub

