GitHub opens ReviewBench, a public benchmark you can run your own AI code reviewer against
The research-preview benchmark scores code review agents on 219 real public PRs across 19 languages, with the full dataset, rubric, and Claude Sonnet 5 judge published, and you can sign in with GitHub, bring a container image plus your own model key, and test on a 25-PR set before a full leaderboard run. GitHub built it to tune Copilot code review, the golden set leans on LLM judging, and leaderboard entries need maintainer approval, so treat the rankings as a starting point.
