Measurement Invariance as an Item-Screening Criterion for Large Language Model Benchmarks: A Position Note
Item response theory is increasingly applied to large language model benchmarks. Recent work uses item response models to select small informative subsets of items, to adapt item administration to model ability, to flag likely label errors, and to estimate latent ability instead of relying on raw accuracy. Recent work...