Improving the Reliability of LLM Evaluation Metrics via Human-in-the-Loop Validation
Abstract
Large language models (LLMs) are increasingly evaluated using automated metrics such as ROUGE, BERTScore, and perplexity. However, these scores often fail to reflect real-world usefulness, particularly for tasks requiring complex reasoning or agentic behavior. This paper examines the risks of misaligned LLM evaluation and identifies key failure modes where metrics reward outputs that are semantically incorrect or provide little value to users. We propose a multi-layer evaluation framework that combines statistical metrics, model-based judges (for example, G-Eval and SelfCheckGPT), and human-in-the-loop (HITL) review. Using stress-tested datasets and perturbation analyses, we expose vulnerabilities in widely used metrics and quantify their false-alignment rates. To mitigate these issues, we introduce a disagreement-driven evaluation loop and a multi-metric aggregation module that prioritizes safety and factuality. Overall, our framework offers practical design principles for building more trustworthy LLM evaluation pipelines and for advancing task-aligned, semantics-focused measurement standards.