Skip to content
Conference

Discovering and Repairing Blind Spots in LLM-as-a-Judge Evaluation of Multi-Turn AI Systems

Jul 2026 · 2026 International Conference on Intelligent and Sustainable AI Systems (ICOSAAS) · pp. 1611-1618 · 0 citations · 17 references

Abstract

The widespread use of Large Language Models (LLMs) as automated evaluators for multi-turn conversational AI systems is due to their scalability and adaptability. Nonetheless, LLM-as-a-judge systems often have systemic blind spots, leading to the neglect of certain flaws because to insufficient, inflexible, or biased assessment standards. This project investigates if disparities between human annotations and LLM evaluations might function as a diagnostic instrument to identify missing evaluation criteria and improve rubric design. By developing a rubric, we provide a systematic framework that may identify indications of disagreement, categorize them to uncover concealed defect patterns, and initiate new aspects of evaluation. The proposed improvements to the rubric are assessed using previously inspected datasets to verify its universality and reliability. We provide a confidence-aware abstention mechanism that detects doubtful predictions and refers them to human assessors to enhance safety and dependability. Experimental results demonstrate that the proposed technique improves defect memory, optimizes the calibration of confidence and accuracy in defect prediction, decreases false alarms, and lowers total annotation costs. Our findings suggest that human disagreement should be regarded not just as noise, but as a crucial signal that might enhance LLM-based assessment systems over time.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.