Dependence-Aware Aggregation Improves LLM-as-a-Judge Accuracy by 9% to 14%
Original titleWhen LLM judges agree, should we believe them?
AISummary
Amazon researchers proposed a dependence-aware method for aggregating LLM judges' votes, using an Ising model to account for correlated errors among judges.
The approach outperformed a weighted majority-vote baseline by 9% to 14% on standard metrics across three binary tasks, including relevance classification, where it reached 0.912 accuracy versus 0.820.
The method is unsupervised, learning from judge outputs without human reference labels.
Source: Amazon Science · amazon.sciencePublished · added here