Discovering and Repairing Blind Spots in LLM-as-a-Judge Evaluation of Multi-Turn AI Systems
Abstract
The widespread use of Large Language Models (LLMs) as automated evaluators for multi-turn conversational AI systems is due to their scalability and adaptability. Nonetheless, LLM-as-a-judge systems often have systemic blind spots, leading to the neglect of certain flaws because to insufficient, inflexible, or biased assessment standards. This project investigates if disparities between human annotations and LLM evaluations might function as a diagnostic instrument to identify missing evaluation criteria and improve rubric design. By developing a rubric, we provide a systematic framework that may identify indications of disagreement, categorize them to uncover concealed defect patterns, and initiate new aspects of evaluation. The proposed improvements to the rubric are assessed using previously inspected datasets to verify its universality and reliability. We provide a confidence-aware abstention mechanism that detects doubtful predictions and refers them to human assessors to enhance safety and dependability. Experimental results demonstrate that the proposed technique improves defect memory, optimizes the calibration of confidence and accuracy in defect prediction, decreases false alarms, and lowers total annotation costs. Our findings suggest that human disagreement should be regarded not just as noise, but as a crucial signal that might enhance LLM-based assessment systems over time.