Beyond Imitation: A Framework and Benchmark for LLM-Assisted Peer Review
Peer review needs support, not substitutes. Accepted at TMLR: our new paper on using AI to help reviewers catch errors in research papers.
https://arxiv.org/abs/2610.11087
As research submissions grow, so does the workload for the experts who evaluate them. AI review systems are emerging as a potential solution, but much of their development and evaluation focuses on how closely they imitate human reviews.
In our work “Beyond Imitation: A Framework and Benchmark for LLM Assisted Peer Review”, we explore how AI can support the review process without losing the human touch. We focus on one demanding but essential task: catching errors in research papers.
We build an automated pipeline that deliberately introduces contradictions into papers by adding statements that conflict with information elsewhere in the manuscript. These planted errors give us clear targets for testing whether AI reviewers can spot and explain what is wrong.
Not all errors carry the same weight. Some undermine a paper’s central findings; others affect smaller details. We map the connections between each paper’s claims, methods, and evidence in a knowledge graph to estimate the severity of each contradiction and better assess what different systems can catch.
We also introduce Multi-Layered Review, an AI review system inspired by the Three-Pass Approach to reading research papers. It first outlines the main ideas, then examines the details and potential weaknesses, and finally brings its observations together into a review. The idea is simple: understand the paper before judging it.
In our evaluations, the system detected more errors than the other review systems we tested, including on papers withdrawn because of real mistakes. Its feedback emphasized different aspects of the work from human reviews, offering a complementary perspective, while its assessments of paper quality remained broadly consistent with human judgments.
Our goal is to give reviewers useful support in checking research, with human expertise and judgment at the center.
Open Review: https://openreview.net/forum?id=7iX2Z2bPFB
