Skip to content
Review

JudgeStealer: Extracting LLM Judging Capabilities across Evaluation Protocols

Aug 2026 · 0 citations · 48 references
Computer Science

TL;DR

JUDGESTEALER is proposed, the first query-efficient model extraction framework for replicating judging capabilities across pointwise scoring, pairwise comparison, and listwise ranking protocols and demonstrates robustness against representative extraction defenses.

Abstract

Large language model (LLM) judges are increasingly used across various evaluation scenarios, making their judgment capabilities valuable intellectual property. However, black-box access exposes these capabilities to model extraction attacks. Existing extraction methods do not specifically target LLM judges and provide limited support for multiple evaluation protocols under restricted query budgets. In this study, we propose JUDGESTEALER, the first query-efficient model extraction framework for replicating judging capabilities across pointwise scoring, pairwise comparison, and listwise ranking protocols. JUDGESTEALER exploits the strong cross-protocol agreement to acquire pointwise scores and transform them into pairwise and listwise supervisions without additional victim queries. To capture informative judge patterns and improve query efficiency, JUDGESTEALER dynamically selects pointwise inputs based on semantic diversity, predictive uncertainty, and potential judge biases. It further applies score smoothing and multi-protocol review to preserve the ordinal structure of scores and mitigate catastrophic forgetting during surrogate adaptation. Extensive experiments on state-of-the-art LLM-as-a-judge and reward models show that JUDGESTEALER consistently outperforms existing extraction baselines, achieving up to 73.3%, 87.0%, and 71.6% accuracy for pointwise, pairwise, and listwise evaluation, respectively. JUDGESTEALER also remains effective across different sur- rogate model scales, adaptation strategies, and reasoning settings. Moreover, JUDGESTEALER demonstrates robustness against representative extraction defenses.

View source

Similar papers

#machine learning Preprint Sep 2026

Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation

This work investigates LLM-based evaluators of natural language generation quality mechanistically through an eight-attack perturbation taxonomy across the Readability and Adequacy dimensions of NLG quality, a generation pipeline that produces paired clean and corrupt summaries with controlled error intensity and expli...

Himil Vasava, Ming-Zhou Jiang · 0 citations
Preprint Aug 2026

Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation Independence

It is shown that prior scores, even when included only as context metadata, anchor judgments and systematically shift ratings toward their values, and effective mitigation must be validated for the intended model and task or domain.

A. Kapetanović, Kemal Altwlkany, Andro Merćep et al. · 0 citations
Conference Aug 2026

Improving the Reliability of LLM Evaluation Metrics via Human-in-the-Loop Validation

Large language models (LLMs) are increasingly evaluated using automated metrics such as ROUGE, BERTScore, and perplexity. However, these scores often fail to reflect real-world usefulness, particularly for tasks requiring complex reasoning or agentic behavior. This paper examines the risks of misaligned LLM evaluation...

Karthik Babu Manam, Vincent Koc, Jamshaid Iqbal Janjua · 0 citations
#small language model Preprint Aug 2026

Localize-Then-Decide Guarantees for LLM Judgments

This work proposes a Localize-Then-Decide framework, which restores the monotonic relationship between confidence and disagreement risk and enables high-probability agreement guarantees in large language models.

Xinyue Li, Yi Zhou, Guanqun Cao et al. · 0 citations
Conference Jul 2026

Discovering and Repairing Blind Spots in LLM-as-a-Judge Evaluation of Multi-Turn AI Systems

The widespread use of Large Language Models (LLMs) as automated evaluators for multi-turn conversational AI systems is due to their scalability and adaptability. Nonetheless, LLM-as-a-judge systems often have systemic blind spots, leading to the neglect of certain flaws because to insufficient, inflexible, or biased as...

Vinay Gummadavelli, Usman Imtiaz, Prabagaran A et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.