Skip to content

Author

Saanvika Reddy Kondapalli

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Jul 2026

Discovering and Repairing Blind Spots in LLM-as-a-Judge Evaluation of Multi-Turn AI Systems

The widespread use of Large Language Models (LLMs) as automated evaluators for multi-turn conversational AI systems is due to their scalability and adaptability. Nonetheless, LLM-as-a-judge systems often have systemic blind spots, leading to the neglect of certain flaws because to insufficient, inflexible, or biased assessment standards. This project investigates if disparities between human annotations and LLM evaluations might function as a diagnostic instrument to identify missing evaluation criteria and improve rubric design. By developing a rubric, we provide a systematic framework that may identify indications of disagreement, categorize them to uncover concealed defect patterns, and initiate new aspects of evaluation. The proposed improvements to the rubric are assessed using previously inspected datasets to verify its universality and reliability. We provide a confidence-aware abstention mechanism that detects doubtful predictions and refers them to human assessors to enhance safety and dependability. Experimental results demonstrate that the proposed technique improves defect memory, optimizes the calibration of confidence and accuracy in defect prediction, decreases false alarms, and lowers total annotation costs. Our findings suggest that human disagreement should be regarded not just as noise, but as a crucial signal that might enhance LLM-based assessment systems over time.

Vinay Gummadavelli, Usman Imtiaz, Prabagaran A et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.