Qualification by Calibration: A Readable Benchmark for Admitting Language Models to Human-Computation Tasks
Evaluating a diverse panel of contemporary models under three administration protocols, it is shown that a naive harness, with a fixed token budget and an unaudited parser, manufactures failing workers out of competent ones, misreading a model that answers essentially every item correctly as badly inaccurate, answer-biased, and overconfident.