Skip to content
Read the original: Ai2 (Allen Institute for AI)·Published AI score56/100

Ai2 introduces BenchMIRT to audit what individual LLM benchmark questions measure

Original titleBenchMIRT: What are LLM benchmarks actually measuring?

AISummary

Ai2 introduces BenchMIRT, a multidimensional item response theory method that audits LLM benchmarks at the level of individual prompts.

Trained on results from 100 LLMs across 16 benchmarks, it recovered safety and general reasoning as the two dominant dimensions, and found BBQ aligns more with general reasoning than safety.

Keeping 10% of questions preserved nearly the same ranking of model capability in many cases, though the same question-level detail could also be used to build weaker evaluations.

Read the original allenai.org

Source: Ai2 (Allen Institute for AI) · allenai.org