Skip to content
Preprint

A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation

Aug 2026 · 0 citations · 57 references
Computer Science

TL;DR

The results show that high judge agreement can coexist with weak sensitivity to changes in the construct being evaluated, motivating joint reporting of invariance and sensitivity and auditing the validation set itself.

Abstract

LLM-as-a-judge evaluation is usually assessed by agreement and robustness to surface perturbations, but reliability does not establish construct validity. We formalize construct validity for an evaluator as a two-dimensional profile: invariance S, the probability that a verdict is unchanged under construct-preserving edits, and construct sensitivity R, the probability that it changes under minimal construct-changing edits. We show that S and R are independent and that no scalar summary preserves all relevant comparisons. We measure the profile across 7 judges and 4 domains using 7 construct-changing intervention types and 5 register-only controls, with intervention direction determined by human annotators and generation, verification, and judging assigned to disjoint model families. At matched invariance S>= 0.90, judges average S = 0.945 but R = 0.319. Sensitivity also differs between scope and strength edits: R_scope = 0.383 versus R_strength = 0.262, a +0.121 gap with the same sign for all 7 judges. We further audit five public label sets and find that surface-only predictors reproduce 55%-67% of labels in paired mode, including 67.4% of MT-Bench human votes. These results show that high judge agreement can coexist with weak sensitivity to changes in the construct being evaluated, motivating joint reporting of invariance and sensitivity and auditing the validation set itself.

View source

Similar papers

Preprint Aug 2026

Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation Independence

It is shown that prior scores, even when included only as context metadata, anchor judgments and systematically shift ratings toward their values, and effective mitigation must be validated for the intended model and task or domain.

A. Kapetanović, Kemal Altwlkany, Andro Merćep et al. · 0 citations
Jul 2026

Which Values Do LLMs Confuse? A Schwartz-Based Recognition Study

This work evaluates 21 instruction-tuned LLM runs under a fixed ranked-response protocol, showing that models often locate the correct motivational region while ranking close alternatives unstably, and motivates value-recognition evaluation that combines exact accuracy, ranked recovery, and directed error analysis.

A. Chetvergov, S. Ukolov, Timofei Sivoraksha et al. · 0 citations
Jul 2026

LLM Judges Can Be Too Generous When There Is No Reference Answer

LLM judges are increasingly being used to evaluate open-ended model responses, often in no-reference settings where a ground-truth answer is unavailable. However, can they reliably assess in such evaluation setups? We explore this question in this paper through a two stage pipeline with a) calibration experiments that assess the judge model's knowledge of the task it is evaluating, and b) sensitivity experiments that assess how the judge model's performance is impacted by the presence and positioning of the reference answer in the prompt. Across experiments covering three languages, we show that the judge models we evaluated tend to over-credit incorrect answers in the absence of a reference answer, and adding reference answer information to the prompt flips the judge model's correct/incorrect decisions by as much as 85% in some experimental settings. Comparison with a subset of human annotations shows that these reference-driven changes generally align with human judgments. Our results emphasize the need for calibrating the LLM judges with a sample with reference-aware evaluation before using them in reference-free setups reliably, and our methodology provides a blueprint for researchers and practitioners in doing such calibration of LLM judges for other tasks.

Chalamalasetti Kranti, Sowmya Vajjala · 1 citation
#artificial intelligence Preprint Aug 2026

Validity-Aware Jailbreak Evaluation for Large Language Models

Jailbreak robustness has become central to large language model (LLM) safety evaluation, yet prevailing methodologies rely primarily on refusal behavior, semantic resemblance, and intent-matching heuristics that emphasize linguistic plausibility rather than correctness. We identify a key limitation in existing evaluations: many jailbreak intents depend on instructional validity rather than epistemic factuality, allowing realistic-looking responses to be labeled successful despite being factually or procedurally incorrect. To address this gap, we propose Sequential Epistemic and Action-Level Validation (SEAV), a verification-centric jailbreak evaluation framework that decomposes responses into ordered steps and evaluates both validity and correctness. SEAV combines LLM-as-a-judge mechanisms for semantic interpretation with retrieval-grounded verification using external knowledge sources, assessing whether generated content is factually correct, structurally consistent, and operationally capable of advancing harmful objectives. Empirically, SEAV cuts the false-positive rate on SD-A (a curated strategic-dishonesty diagnostic) by 14.9\,pp vs. the strongest baseline, and reclassifies 22.1\%--51.0\% of sampled prior-labeled successes as invalid across three of four public benchmarks. Together, these results show that enforcing correctness substantially reshapes measured robustness: many previously labeled jailbreak successes are reclassified as invalid, and results are stable across the tested search backends and evaluator models. Code and data are available at https://github.com/Ardor-Wu/SEAV.

Qilong Wu, Sahil Wadhwa, Pranab Mohanty et al. · 0 citations
Open access Aug 2026

ConsensusGrade: A Human-Variability-Aware Framework for Evaluating LLM-Based Automated Grading

Background: Evaluation of LLM-based automated grading often relies on comparison with a single human score, which can obscure meaningful variability among raters of open-ended answers. This study introduces ConsensusGrade, a consensus-aware framework that treats the human reference as a scoring envelope rather than as a single point. Methods: We analyzed 1000 open-ended student answers from 100 students across 10 questions, each graded by four evaluators. Six previously generated and aligned automated grading configurations from GradeAgentOps were compared with the four-rater human reference. The score sets were generated using Llama 3.3 70B Instruct as the primary grader, with Qwen 2.5 14B Instruct for semantic repair. Results: Human evaluators showed meaningful agreement, with ICC(A,1) = 0.712, but exact four-rater agreement occurred in only 2.2% of records. Broad score dispersion occurred in 59.0%. All automated configurations showed negative bias relative to the human median. FULL achieved 68.5% inside-envelope positioning and a chance-adjusted score of 0.454; under the central-trimmed envelope, this rate decreased to 34.3%, while configuration ordering was preserved. Conclusions: ConsensusGrade provides a diagnostic framework for interpreting automated scores relative to observed human variability; inside-envelope rates should not be interpreted as stand-alone measures of grading accuracy.

Cătălin Anghel, A. Anghel, Mihai Vlase et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.