Skip to content
Preprint

When Trust Meets Truth: Trust-Truth Separability in LLM-as-Judge

Aug 2026 · 0 citations · 29 references
Computer Science

TL;DR

Results show that current LLM-as-Judge protocols should not treat trust scores as independent evidence for truth judgments, suggesting weaker separations between trust and truth judgment.

Abstract

LLM-as-Judge systems can produce multi-dimensional evaluations, such as trustworthiness, reliability, and factuality, and these outputs are often interpreted as independent evidence. We test this assumption for a common pair of judgments: trust scoring and binary truth classification. On correctness-controlled QA, LLM judges align trust scores with truth verdicts more tightly than human behavioral reference, suggesting weaker separations between trust and truth judgment. We then apply stress tests by changing only source cues of identical QA between Human and AI. Source attribution shifts not only trust scores but also truth verdicts and logit-derived correct-side probabilities. Results show that current LLM-as-Judge protocols should not treat trust scores as independent evidence for truth judgments.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Can We Trust LLM Judges: A Study of Capability-Dependent Biases and Multi-Judge Ensemble for Bias Calibration

LLMs are increasingly used as automated judges for model training and evaluation, yet individual judges exhibit systematic biases that undermine reliability. Much of prior work has studied biases in pairwise LLM-as-a-judge settings; in this paper, we focus on absolute scoring tasks, which mirror more realistic use case...

Gemma Zhang, Prachi Badarayani, Asmi Kumar et al. · 0 citations
#machine learning Preprint Sep 2026

GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents

Comparing and selecting task-oriented LLM agents increasingly relies on a low-cost offline evaluation gate: persona-driven LLM user-simulators converse with each candidate, an LLM-as-a-judge scores the transcripts, and the higher-scoring agent is promoted. We introduce GAUGE, a reusable offline protocol that measures w...

Umesh Bodhwani, Thanh Tran, Kai-Lin Wei · 2 citations
#artificial intelligence Preprint Sep 2026

Who Judges Matters: Measuring Family-Conditioned Preference in LLM-as-Judge Panels

Who the judge is can affect an LLM-as-judge result, but measuring that effect without confusing it with candidate quality is difficult. We study four open-weight families (Llama 3.1, Qwen 2.5, Gemma 2, and Yi 1.5) in a fully crossed pairwise design with 9,312 judgments. A common per-family statistic is strongly confoun...

David Ababio Awuni, Luke Achenie, Benjamin Tei Partey et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus

Consensus among LLM judges is often taken as strong evidence that a decision is correct. This assumes that judges make their errors independently. In practice, LLM judges are often trained and evaluated in similar ways, so they can make the same mistakes. We study how this dependency affects the reliability of consensu...

Elias Hossain, Niloofar Yousefi, Ser-Nam Lim · 0 citations
#artificial intelligence Preprint Sep 2026

Accounting for Bias Enables Sustainable LLM Evaluation

LLM-as-a-judge has become the de facto standard for scalable, subjective evaluation, yet current leaderboards compensate for systematic measurement bias by running ever more comparisons, an approach that is both statistically unsound and computationally wasteful. The root cause is an incomplete measurement model, treat...

H. Katoch, D. A. Selby, Gerrit Großmann et al. · 0 citations
#artificial intelligence Preprint Sep 2026

How to Run Statistics over LLM Judges and Trust the Results: Calibrated Inference for Small-Sample AI Evaluation with evalstats

Researchers across academia increasingly base significance claims on LLM judge scores and small-sample AI evaluations. Yet without well-calibrated confidence intervals (CIs), hypothesis tests, and judge-bias corrections, such claims are unreliable. We address these issues in several contributions. First, we find that r...

Ian Arawjo · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.