Position: Toward a Metric Typology for Language Model Evaluation
Abstract
The critique of scalar benchmark rankings as proxies for model quality is now well-established (Raji et al., 2021; Wallach et al., 2025; Bean et al., 2025; Gehrmann et al., 2021). What the field still lacks is a shared structural vocabulary for comparing, combining, and contextualizing metric design choices. This paper provides that vocabulary: a four-primitive typology—representation ( ϕ ), comparison ( D ), aggregation ( A ), and context ( C )—under which existing metrics (BLEU, BERTScore, nDCG, LLM-as-judge, calibration scores, agentic outcome measures) are explicit parameterizations of a common form. This typology is paired with a measurement–decision split: metrics are noisy estimators of latent constructs, and model selection is context-dependent Pareto optimization over construct estimates, not over raw scores. The typology makes implicit metric assumptions comparable and debatable rather than hidden inside a single number.