Skip to content

When Rubrics Fail: Hallucinations Reveal Blind Spots in Medical AI Evaluation

Sep 2026 · 0 citations
Computer Science

TL;DR

A taxonomy of medical hallucination types and a clinician-validated error-injection pipeline that creates matched correct and error-injected responses are developed, finding that more specific rubrics better distinguish correct from hallucinated responses.

Abstract

Hallucinations can undermine clinician trust in LLMs, making it important that evaluation methods capture clinically relevant errors. Rubric-based evaluation has become the leading approach for assessing LLMs in medicine, but it is unclear whether rubric scores reflect such errors. We first study this in a controlled setting using MedHallu, finding that more specific rubrics better distinguish correct from hallucinated responses. To test this systematically, we develop a taxonomy of medical hallucination types and a clinician-validated error-injection pipeline that creates matched correct and error-injected responses. Across HealthBench, HealthBench Professional, and LiveMedBench, our clinically relevant hallucinations are missed by rubrics, often leaving scores unchanged. We find that rubrics are most effective when explicitly checking facts, and are less effective for additional or unexpected errors they do not anticipate. A preliminary retrieval-based factuality check recovers some of the rubric-blind errors, suggesting a complementary approach. These findings reveal systematic blind spots in current medical evaluation of LLMs and suggest that rubric scores alone are insufficient to establish clinical reliability, potentially undermining clinician trust and confidence in clinical deployment.

View source

Similar papers

Review Open access Sep 2026

Retrieve-Then-Verify for Evaluating Evidence Support and Hallucination in Large Language Model–Generated Medical Information: Empirical Study

Retrieval-based evidence verification provides a reproducible and transparent approach for evaluating the reliability of AI-generated medical information, with direct relevance to digital health practice, evidence-based medicine, and medical informatics.

Zhao-Hui Liang, Cynthia Sheffield, Gisela Butera et al. · 0 citations
Conference Open access Sep 2026

Factual Hallucination in Medical Large Language Models: Typology, Evaluation, and Systematic Governance

A fine-grained hallucination classification framework based on medical knowledge graphs is proposed and a three-tier risk stratification scheme that references FDA medical AI software risk classification standards is established, highlighting the unique characteristics of the medical domain.

Zi-Jiao Liu · 0 citations

Right for the Wrong Reasons: A Benchmark for Hallucination and Clinical Safety in AI Health Triage

An empirical benchmark for evaluating clinical triage systems that assesses explanation quality alongside decision outcomes, and provides a reproducible, checkpoint-based evaluation pipeline and outline a roadmap for bias stress-testing, hallucination mitigation, and open benchmark release.

S. Marimuthu, Patricia L. Mabry, HealthPartners.Com · 0 citations
#artificial intelligence Preprint Sep 2026

Hallucination Neurons and Where to Find Them: An Investigation into the existence of Hallucination Neurons

Interpretable machine learning for Large Language Models (LLMs) increasingly relies on sparse probing methods that identify small sets of neurons claimed to detect and causally influence behaviors such as factuality recall, safety alignment, and hallucination. These claims have important implications for model auditing...

Huseyin Cavus, Sebin Sabu, J. Spear et al. · 0 citations
#natural language process... Preprint Oct 2026

Single-Pass Uncertainty Heads for Claim-Level Hallucination Detection in Persian Medical Language Models

Hallucination detection is particularly important for medical language models, but repeated-sampling approaches are expensive and existing uncertainty-head resources do not directly transfer to a new backbone and language. We adapt the LLM Uncertainty Head (LUH) framework to Aya-Expanse-8B-based Persian medical models,...

Mehrdad Ghassabi, Pedram Rostami, Hamidreza Baradaran Kashani et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.