2026· European Journal of Prosthodontics and Restorative Dentistry· Vol 34· 0 citations
TL;DR
A narrative review distinguishes fabrication, the invention of an entirely non-existent source, from unfaithfulness, the citation of a genuine source in support of a claim it does not contain, and argues that the two failure modes require different detection strategies.
Abstract
Large language models (LLMs) have introduced a new research integrity threat into the biomedical literature i.e. fabricated references that appear authentic but correspond to no existing publication. This narrative review distinguishes fabrication, the invention of an entirely non-existent source, from unfaithfulness, the citation of a genuine source in support of a claim it does not contain, and argues that the two failure modes require different detection strategies. Early evaluations of ChatGPT-3.5 reported fabrication rates as high as 69% for medical questions, and although newer models perform better, residual rates in the range of 15 to 20% remain in current systems. The review examines five interconnected domains: the autoregressive architecture that makes fabrication structurally likely rather than incidental; the marketing of AI-powered search tools, which frames speed and accuracy as complementary rather than competing, and which disproportionately misleads early-career researchers in resource-limited settings; the compounding effect of AI paraphrasing tools, which can sever the link between a claim and its supporting citation even when no new reference is fabricated; the cascading consequences that extend from manuscript rejection to contamination of the clinical evidence base; and to the current, fragmented state of detection and editorial policy. A recurring finding across the reviewed literature is that professional presentation does not reliably track factual accuracy, so reviewer confidence in AI-assisted text is a poor substitute for verification. The review identifies real-time reference verification integrated into AI-assisted writing workflows as the central unmet need, and proposes a three-tier response spanning researcher-level verification habits, editorial audit infrastructure, and discipline-specific AI literacy training. The findings support the conclusion that fabricated references in medical research reflect a structural property of current language models rather than a transient limitation, and that addressing the problem requires coordinated action rather than reliance on model improvement alone.
A structured survey of AI hallucinations, synthesizing prior research to clarify their evolving definitions, underlying causes, and implications for Information Systems positions AI hallucinations as socio-technical phenomena with direct implications for trust, decision-making, and governance.
Under standardized zero-shot, retrieval-disabled web-interface conditions, LLMs generated substantial numbers of inaccurate and fabricated NCC citations, which should be verified across reliable databases before use in clinical, educational, or scholarly work.
Ali Seifi, A. Seyfi· Critical Care Explorations· 0 citations
This article examines the risk of hallucination when using large language models (LLMs) for financial research and demonstrates the unreliability of off-the-shelf tools such as ChatGPT for tasks requiring precise financial reasoning. A simple, replicable experiment was conducted using the free online version of ChatGPT...
Hongbok Lee· International Journal of Eco...· 0 citations
A fine-grained hallucination classification framework based on medical knowledge graphs is proposed and a three-tier risk stratification scheme that references FDA medical AI software risk classification standards is established, highlighting the unique characteristics of the medical domain.
"AI psychosis"has entered public and clinical discourse as a label for the onset or exacerbation of psychotic symptoms, most commonly delusions, following intensive interaction with large language model (LLM)-based chatbots. Current evidence is limited to media reports, case reports, and early observational data, yet t...
Joshua Au Yeung, H. Morrin, Vincent Ng et al.· 0 citations
An empirical benchmark for evaluating clinical triage systems that assesses explanation quality alongside decision outcomes, and provides a reproducible, checkpoint-based evaluation pipeline and outline a roadmap for bias stress-testing, hallucination mitigation, and open benchmark release.
S. Marimuthu, Patricia L. Mabry, HealthPartners.Com· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.