It is argued that each NLP subfield should pair its main performance metric with a calibration score and call for treating calibration as an essential property of every model rather than a niche topic.
Abstract
Calibration of language models -- the alignment between expressed or implicit confidence and empirical correctness -- is a well-studied subfield within NLP. Methods to measure it already exist. The problem is adoption: outside this subfield, NLP research regularly introduces new models, datasets, and benchmarks without checking whether the model's confidence scores are meaningful. We argue that this adoption gap is a major obstacle to trustworthy LLM evaluation. Miscalibration causes problems in two distinct areas: at deployment, where overconfident mistakes cause real harm, and inside the research pipeline, where methods like LLM-as-a-judge, synthetic data generation, and active learning rely on calibrated confidence without verifying it. Standard calibration metrics only require two inputs per example: a confidence score and a correctness judgment. Most benchmarks in use today already provide both, meaning calibration can be reported immediately. For open-ended generation, however, defining these two inputs is still an open challenge. We argue that each NLP subfield should pair its main performance metric with a calibration score and call for treating calibration as an essential property of every model rather than a niche topic.
A lightweight quality-assessment protocol is presented for LLM-generated synthetic training data and applied to 13,579 synthetic user reviews generated from GitHub issues across four open-source Android applications, high-lighting the need for hybrid human-AI verification when synthetic data is used in security-critica...
This work argues that an LLM judge running in a production system is better understood as having a lifecycle: it must be built, trained, deployed, and continuously maintained as the surrounding data evolves, and each phase poses distinct technical and operational challenges.
Emma Kong, J. Tan, Ishan Gupta et al.· 0 citations
This work investigates LLM-based evaluators of natural language generation quality mechanistically through an eight-attack perturbation taxonomy across the Readability and Adequacy dimensions of NLG quality, a generation pipeline that produces paired clean and corrupt summaries with controlled error intensity and expli...
It is shown that prior scores, even when included only as context metadata, anchor judgments and systematically shift ratings toward their values, and effective mitigation must be validated for the intended model and task or domain.
A. Kapetanović, Kemal Altwlkany, Andro Merćep et al.· 0 citations
Pinocchio is introduced, an external calibrator that estimates the correctness of responses from black-box API models that needs only a single forward pass to generate an uncertainty estimate and requires no access to the target model's logits, weights, or internal states.
K. D. Hayes, Arka Pal, Hao-Song Zhang et al.· 0 citations
CruxBench is introduced, a benchmark that grades LLM-generated questions by their Value of Information (VOI): how much a model-proposed crux updates beliefs about a target forecasting question and finds that VOI correlates highly with independent measures of model capability and captures cruxes'usefulness for answering...
Hui Dai, Li-Na Piao, Nick Merrill et al.· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 24, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.