Skip to content
GitHub Blog · AI & ML· Michelle Zhou·· 3d agoPickAI score63

GitHub releases ReviewBench, an open benchmark for AI code review agents

ReviewBench: An open benchmark for AI code review

AI summary

GitHub has released ReviewBench, an open benchmark for evaluating AI code review agents on 219 public pull requests across 19 languages. The benchmark reports grounded and augmented precision, recall, and F1 metrics, and its dataset, rubric, and judge are publicly available. GitHub says ReviewBench predicted the direction of a Copilot code review ensemble experiment's production results before A/B testing.

Why it matters

The post explains how ReviewBench was built and validated, and reports an offline-to-production comparison that shows how well a benchmark predicts real experiment outcomes.

Read the original github.blog

Source: GitHub Blog · AI & ML · github.blog