Skip to content
Conference

Improving the Reliability of LLM Evaluation Metrics via Human-in-the-Loop Validation

Aug 2026 · 2026 7th International Conference on Big Data Analytics and Practices (IBDAP) · pp. 1-6 · 0 citations · 19 references

Abstract

Large language models (LLMs) are increasingly evaluated using automated metrics such as ROUGE, BERTScore, and perplexity. However, these scores often fail to reflect real-world usefulness, particularly for tasks requiring complex reasoning or agentic behavior. This paper examines the risks of misaligned LLM evaluation and identifies key failure modes where metrics reward outputs that are semantically incorrect or provide little value to users. We propose a multi-layer evaluation framework that combines statistical metrics, model-based judges (for example, G-Eval and SelfCheckGPT), and human-in-the-loop (HITL) review. Using stress-tested datasets and perturbation analyses, we expose vulnerabilities in widely used metrics and quantify their false-alignment rates. To mitigate these issues, we introduce a disagreement-driven evaluation loop and a multi-metric aggregation module that prioritizes safety and factuality. Overall, our framework offers practical design principles for building more trustworthy LLM evaluation pipelines and for advancing task-aligned, semantics-focused measurement standards.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.