Skip to content

MAWILE: Multi-Axis Workbench for Inspecting LLM Evaluators

Sep 2026 · 0 citations · 21 references
Computer Science

TL;DR

MAWILE is introduced, a developer-facing workbench for auditing judge sensitivity across four surfaces: the judge prompt, judge rubric, target-system input, and target-system output, and it audits binary, ordinal, and pairwise judges without requiring gold labels.

Abstract

Large language model (LLM) judges provide a flexible and scalable method for evaluating model and agent outputs, but their verdicts can be sensitive to incidental changes in the evaluated response, judge instructions, and scoring rubric. Existing systems examine important subsets of these failure modes, but auditing a configured judge requires testing both the judge instrument and the items it evaluates. We introduce MAWILE, a developer-facing workbench for auditing judge sensitivity across four surfaces: the judge prompt, judge rubric, target-system input, and target-system output. Given a user-supplied judge and representative evaluation items, MAWILE constructs and validates controlled perturbations, re-executes the judge, and localizes the resulting sensitivity. Each perturbation declares whether the verdict should remain invariant or change in a specified direction, allowing the same system to measure both robustness to irrelevant variations and sensitivity to meaningful changes. MAWILE audits binary, ordinal, and pairwise judges without requiring gold labels. The code for this tool is available at: github.com/megagonlabs/mawile-judge.

View source

Similar papers

Preprint Sep 2026

JEV as a Judge for Agent Trace Security: An Empirical Comparison with Generative LLM Judges

Security evaluation of tool-using agents requires judging actions in context, yet generative judges add latency, explanation overhead, and output-validation failures. We study whether JEV, a typed decision model, offers a useful alternative for retrospective trace classification. We evaluate JEV and four generative jud...

Zhi-Qiang Wang, Yi-Chao Gao · 2 citations · ⚡1
Review Aug 2026

JudgeStealer: Extracting LLM Judging Capabilities across Evaluation Protocols

JUDGESTEALER is proposed, the first query-efficient model extraction framework for replicating judging capabilities across pointwise scoring, pairwise comparison, and listwise ranking protocols and demonstrates robustness against representative extraction defenses.

Chen Chen, Yao-Lin Chen, Xue-Han Sun et al. · 0 citations
Preprint Aug 2026

A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation

The results show that high judge agreement can coexist with weak sensitivity to changes in the construct being evaluated, motivating joint reporting of invariance and sensitivity and auditing the validation set itself.

Jian-Lin Chen, Wen-Hui Chen, Zi-Yao Lin et al. · 4 citations
#artificial intelligence Preprint Aug 2026

Commit-first LLM judging inherits the judge's own errors

commit-first judging does not remove the anchor that gets gamed, it moves it from the candidate to the judge's own answer, so evaluation is only as good as the judge is at the task, so evaluation is only as good as the judge is at the task.

Idil Gozel · 1 citation
#artificial intelligence Preprint Aug 2026

ExecRubrics: Executable Tool-Augmented Rubrics for Verifiable and Efficient Long-Form Evaluation

The results suggest a novel way of approaching automated evaluation, by offering a faster, more explainable, and less ambiguous alternative to black-box rubric evals, particularly in high-stakes domains such as healthcare and banking where precision and auditability are critical.

Kaustubh D. Dhole, Charles L. A. Clarke, E. Agichtein · 0 citations
#artificial intelligence Preprint Sep 2026

Evaluating and Benchmarking the System One Model Jev

Jev answers MMLU's calculation-heavy questions more accurately than other MMLU questions (94% vs. 91%), whereas both open models, and all three on C-Eval, find them harder.

Tobias Deußer, L. Sparrenberg, R. Sifa · 4 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.