Checking the Check: Controls, Corpus Alignment, and Scale Invariance in Evaluation Software
Evaluation frameworks turn model outputs into the scores that benchmarks, leaderboards and deployment decisions rely on, and the scoring code itself also requires direct checks. We test four evaluation frameworks (DeepEval, TruLens, lighteval and HELM) and one metrics library (TorchMetrics) with three checks: does a sc...