Skip to content
Preprint

Human-Anchored Inference for Ranking New Models with Large Language Model Judges

Sep 2026 · 0 citations · 37 references
Mathematics

Abstract

Human pairwise comparisons provide a reference for evaluating large language models (LLMs), but collecting sufficient judgments for each new release is costly and time-consuming. LLM judges offer a scalable alternative, although their comparisons may differ systematically from human preferences and across judges. We study the ranking of a new model that has received LLM-judge comparisons but no human comparisons. We propose ANCHOR (ANchored Comparisons for Human-reference inference with Orthogonal Riesz correction), which uses historical human and LLM comparisons to learn judge-specific sensitivities to human score differences and feature-dependent judge biases. These estimates are then used to infer the new model's human-reference score from its judge comparisons. The framework allows the feature distribution to change between historical and new-model comparisons. For inference, we construct a Neyman-orthogonal estimator through a joint Riesz correction that removes the first-order effects of estimating the historical human scores, judge sensitivities, and bias functions. We establish identification, convergence rates, and asymptotic normality with consistently estimable variance, and show that ANCHOR attains the semiparametric efficiency bound. Simulations demonstrate gains in score estimation and ranking accuracy. On Chatbot Arena, ANCHOR achieves the lowest score RMSE and insertion MAE among competing methods, with narrower score intervals on average.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.