Skip to content
Conference Open access

Factual Hallucination in Medical Large Language Models: Typology, Evaluation, and Systematic Governance

Sep 2026 · Exploring Science Academic Conference Series · 0 citations · 26 references

TL;DR

A fine-grained hallucination classification framework based on medical knowledge graphs is proposed and a three-tier risk stratification scheme that references FDA medical AI software risk classification standards is established, highlighting the unique characteristics of the medical domain.

Abstract

Large language models (LLMs) have demonstrated transformative potential in clinical documentation generation, diagnostic assistance, and patient consultation. However, their tendency toward “hallucination”— generating semantically fluent but factually inconsistent content with established medical knowledge or input context—constitutes a core safety barrier to clinical deployment. This paper systematically reviews the typology, evaluation frameworks, underlying mechanisms, and mitigation strategies for medica l LLM hallucinations. First, we propose a fine-grained hallucination classification framework based on medical knowledge graphs and establish a three-tier risk stratification scheme that references FDA medical AI software risk classification standards. Sec ond, we analyze multi-factor causes from data, model, and inference dimensions, highlighting the unique characteristics of the medical domain. Third, we systematically review current evaluation benchmarks and methodologies, deeply analyzing the limitations of automated metrics and the potential of LLMs as evaluators. Regarding mitigation strategies, we propose a three-stage classification framework—”training-stage internal intervention —inference-stage external constraints —post-processing collaborative verification” —that provides in-depth analysis of each strategy's internal mechanisms and trade- offs. Finally, we discuss challenges in real-world clinical deployment, emphasizing that the “human-in-the- loop” paradigm remains an irreplaceable safety line.

Read PDF

Similar papers

Conference Aug 2026

Hallucination Detection in Large Language Models

The concern over hallucination, a situation where a model generates fluent but factually inaccurate or fabricated information-has grown with the steady development of Large Language Models like GPT and Gemini. Such results have the potential to erode safety, dependability, and trust in AI driven fields such as journali...

Shanya Kumari, P. Bagane, U. Waghmode et al. · 1 citation
Review Open access Sep 2026

Retrieve-Then-Verify for Evaluating Evidence Support and Hallucination in Large Language Model–Generated Medical Information: Empirical Study

Retrieval-based evidence verification provides a reproducible and transparent approach for evaluating the reliability of AI-generated medical information, with direct relevance to digital health practice, evidence-based medicine, and medical informatics.

Zhao-Hui Liang, Cynthia Sheffield, Gisela Butera et al. · 0 citations
Open access Sep 2026

Errors, Hallucinations, and Clinical Impact of General-Purpose Multimodal Large Language Models in Histopathology

BackgroundGeneral-purpose large language models (LLMs) are increasingly evaluated in diagnostic pathology, but prior studies have largely emphasized diagnostic accuracy rather than how models fail. We evaluated four LLMs for diagnostic performance, pathology-relevant errors and hallucinations, their burden, and potenti...

K. Lami, S. Agarwal, A. Asaturova et al. · 0 citations
Review Open access Sep 2026

Performance and Hallucination Analysis of Large Language Models on European Anesthesiology Examinations: Cross-Sectional Comparative Study

Current-generation LLMs demonstrated consistently high performance across multiple European anesthesiology examinations but continue to produce clinically relevant hallucinations, supporting their role as supervised educational tools rather than autonomous learning resources.

Ștefan Andrei, Thibault Giet, Alexis Belouard et al. · 0 citations
Review Open access 2026

AI Hallucination And Fabricated References: A Growing Crisis For Medical Researchers A Narrative Review

A narrative review distinguishes fabrication, the invention of an entirely non-existent source, from unfaithfulness, the citation of a genuine source in support of a claim it does not contain, and argues that the two failure modes require different detection strategies.

M. Rana, M. Alkhlewi, Turki Abdulaziz Alsohaibani et al. · 0 citations
Review

Survey of AI Hallucinations and Mitigation Survey of AI Hallucinations and Mitigation

A structured survey of AI hallucinations, synthesizing prior research to clarify their evolving definitions, underlying causes, and implications for Information Systems positions AI hallucinations as socio-technical phenomena with direct implications for trust, decision-making, and governance.

R. Thompson, Navid Hashemi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.