Human pairwise comparisons provide a reference for evaluating large language models (LLMs), but collecting sufficient judgments for each new release is costly and time-consuming. LLM judges offer a scalable alternative, although their comparisons may differ systematically from human preferences and across judges. We st...
Xin Zhou, Si-Nian Zhang, Zhan-Yan Yang et al.· 0 citations
This work introduces Counterfactual Harness Search and Evolution (CHASE), which casts harness evolution as constraint generation over valid counterfactual benchmarks and formalizes an ideal shortcut-neutralized benchmark $B_0$ and establishes theoretical guarantees linking finite counterfactual archives to $B_0$ and ch...
Guo-Jun Zhu, Xu Huang, Peng Yin et al.· 0 citations
This work empirically evaluates the primary performance-profile estimator on MATH-500, ScienceQA, MMLU, WinoGrande, HellaSwag, TruthfulQA, GSM8K, and ARC and proves calibration-free identification, unbiasedness for grouped profiles, local oracle optimality within centered linear augmentations, and first-order adaptivit...
Zhi Zhang, Lingfeng Lyu, Yue Kang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.