#natural language process...
May 2026
Auditing LLM Benchmarks with Item Response Theory
An Item Response Theory-based indicator is introduced that surfaces likely mislabels at 95% precision in the top 200 examples across seven preference and multiple-choice benchmarks using responses from 114 models, outperforming a supervised classifier.
Sander Land, D. Bikel
· arXiv.org · 2 citations