#natural language process...
Jun 2026
Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation
This work introduces LongJudgeBench, a comprehensive benchmark for evaluating LLM judges on long-form outputs across diverse real-world scenarios and judging protocols, and systematically evaluates a broad range of LLM judges, covering multiple base models and judging settings.
Junjie Chen, Yuxin Dong, Haitao Li et al.
· arXiv.org · 0 citations