Trustworthy Method Comparison with AI Judges: Estimation and Design under Order, Batch, and Aggregation Effects
Large language models (LLMs) are increasingly used as judges for automated AI evaluation. A common practice is to randomize prompt sequences and average the resulting scores, but its statistical validity remains unclear. We show that LLM evaluation mechanisms can be approximated by a class of Markov generalized linear...