Skip to content
Review

Hallucinations in large language models within academic contexts: a systematic review

Jul 2026 · Journal of Science and Technology Policy Management · pp. 1-28 · 0 citations · 63 references

TL;DR

The findings indicate that hallucinations systematically compromise academic writing quality, distort assessment processes and undermine epistemic trust in scholarly outputs, leading to an integrated conceptual perspective linking hallucinations to epistemic risk, information integrity and digital trust.

Abstract

This study aims to examine the phenomenon of hallucinations in large language models (LLMs) within academic contexts, focusing on their manifestations, causes and implications for academic integrity, research quality and responsible artificial intelligence adoption in higher education. A systematic literature review was conducted in accordance with PRISMA 2020 guidelines. Searches across Scopus, Web of Science and Emerald Insight databases using keywords related to AI hallucination and academic applications, of which 25 peer-reviewed journal articles met the inclusion criteria. Qualitative thematic analysis was performed using NVivo 14 to synthesise evidence on hallucination types, academic applications, impacts and mitigation strategies. Six recurring types of hallucinations were identified, with fabricated or inaccurate citations emerging as the most prevalent. The findings indicate that hallucinations systematically compromise academic writing quality, distort assessment processes and undermine epistemic trust in scholarly outputs. Variation in hallucination rates across models and disciplines highlights their context-dependent nature. Key contributing factors include probabilistic text generation, limitations in training data, insufficient contextual understanding and the absence of robust verification mechanisms. It further contributes a structured classification of hallucination types and a multi-layered governance approach to inform institutional policy and responsible AI adoption. Addressing hallucinations in academic knowledge production is essential for preserving public trust in higher education and safeguarding the societal value of scholarly research. This study advances existing knowledge by developing an integrated conceptual perspective linking hallucinations to epistemic risk, information integrity and digital trust.

View source

Similar papers

Conference Open access 2026

AI Hallucinations in Academic Writing and Addresses

The generative artificial intelligence (AI) has been widely used in academic writing, and its hallucination issue becomes problematic when considering the accuracy and credibility of academic writing. The systematic literature review methodology is used in this paper, in line with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) process to filter articles related to the topic since 2022, with the aim of examining the influence of AI hallucinations on academic writing. The study reveals that AI hallucinations have three kinds of manifestations, which include content distortion, evidence failure and improper argumentation, which comprises factual error, false or irrelevant citation, and superficially coherent arguments lacking evidence. They may destroy the authenticity, transparency, and logic of academic writing. Even if technologies like Retrieval-Augmented Generation (RAG) and Meta-RAG were able to reduce the hallucination effect at least partially, it would not be possible to completely get rid of it. AI need to be considered as an additional tool, and people are still responsible to verify the information and adhere to academic standards. It is crucial to note that developing critical use skills of AI-generated texts by students in educational environments is very important.

Yichen Liu · 0 citations
Review Open access Aug 2026

Hallucination Rate of Peer-Reviewed Citations Generated by Large Language Models in Neurocritical Care

IMPORTANCE: Large language models (LLMs) are increasingly used for scientific literature retrieval, yet their citation accuracy in specialized clinical domains remains poorly characterized. In neurocritical care (NCC), fabricated or inaccurate citations may be difficult to detect without deliberate verification. OBJECTIVES: To evaluate hallucination and fabrication rates of peer-reviewed citations generated by three LLMs across core NCC topics, under constrained zero-shot, memory-only conditions. DESIGN, SETTING, AND PARTICIPANTS: In this cross-sectional, blinded technology performance evaluation, Generative Pretrained Transformer (GPT)-5.3, DeepSeek-V3, and Grok-4 were queried on March 10, 2026, under identical zero-shot, retrieval-disabled web-interface conditions. Ten NCC topics were submitted to each model, and each model generated 10 references per topic, yielding 300 references. MAIN OUTCOMES AND MEASURES: Two NCC experts, blinded to model identity, independently verified each reference against PubMed, DOI, Google Scholar, and CrossRef and scored accuracy using a Hallucination Scale (0–3). The primary outcome was any hallucination, defined as any citation inaccuracy. The secondary outcome was fabrication, defined as a nonexisting complete bibliographic entity. RESULTS: Inter-rater agreement was excellent (κ = 0.91; 95% CI, 0.86–0.96). Overall, 165 of 300 references (55.0%) contained a citation inaccuracy, and 85 of 300 (28.3%) were completely fabricated. DeepSeek-V3 had the lowest hallucination rate (23%; fabrication 8%), followed by GPT-5.3 (69%; fabrication 27%) and Grok-4 (73%; fabrication 50%). Compared with DeepSeek-V3, Grok-4 was 3.17 times more likely to hallucinate (95% CI, 2.03–4.96; p < 0.001), and GPT-5.3 was 3.00 times more likely to hallucinate (95% CI, 1.94–4.63; p < 0.001). Topic-level findings were exploratory and should be interpreted cautiously. CONCLUSIONS AND RELEVANCE: Under standardized zero-shot, retrieval-disabled web-interface conditions, LLMs generated substantial numbers of inaccurate and fabricated NCC citations. Because fabricated references can appear complete and credible, artificial intelligence-generated citations should be verified across reliable databases before use in clinical, educational, or scholarly work.

Ali Seifi, A. Seyfi · 0 citations
Sep 2026

A Q Methodology Study: Exploring Attitudes Toward “AI Hallucinations” in Human–AI Collaborative Academic Writing in Higher Education

While generative artificial intelligence (GenAI) is increasingly integrated into higher education academic writing, it introduces profound epistemic challenges such as AI hallucinations, including fabricated citations and data. Current literature predominantly focuses on macro-level technology acceptance, leaving students’ micro-level subjective perspectives relatively underexplored. This study investigates the subjective perspectives and evaluative orientations of university students when navigating AI hallucinations during human–AI collaborative writing. Employing Q methodology structurally anchored in the Cognitive Process Theory of Writing, 45 humanities and social science students prioritized a refined set of 40 statements. Quantitative factor analysis complemented by qualitative post-sort interviews identified five factor-defined perspectives: Autonomous Directors, Pragmatic Operators, Epistemic Sentinels, Conflicted Compromisers, and Creative Opportunists. These perspectives reflect shared configurations of epistemic trust, authorship control, and verification responsibility, which are interpreted in relation to cognitive agency rather than treated as separately measured constructs or fixed categories of students. The findings offer empirical material for formulating tentative, data-informed pedagogical hypotheses concerning different patterns of epistemic and cognitive support in AI-assisted academic writing; these hypotheses require validation through future intervention-based research.

Unknown authors · 0 citations
Review

Survey of AI Hallucinations and Mitigation Survey of AI Hallucinations and Mitigation

A structured survey of AI hallucinations, synthesizing prior research to clarify their evolving definitions, underlying causes, and implications for Information Systems positions AI hallucinations as socio-technical phenomena with direct implications for trust, decision-making, and governance.

R. Thompson, Navid Hashemi · 0 citations
#generative ai Review Open access Sep 2026

Reliability Risks of Generative AI in Education: A Systematic Review of Hallucinations, Misinformation, Overreliance, and Assessment Validity

Generative AI (GenAI) has introduced reliability risks in education, including hallucinations, fabricated citations, factual inaccuracies, overreliance, and threats to assessment validity. Empirical evidence on these risks has not been systematically synthesized. This review synthesizes that evidence and its consequences for learning and assessment across four research questions. Following PRISMA 2020, Web of Science and Scopus were searched. After duplicate removal, 155 records were screened by two independent doctoral researchers (Cohen's κ = .847). Thirty-five empirical studies were included; 32 (91.4%) were rated high quality using the Mixed Methods Appraisal Tool (MMAT). Thematic synthesis identified four themes: (1) hallucinations and factual inaccuracies taken up as misinformation, moderated by domain expertise; (2) reliability limitations in automated essay scoring, classroom observation, and AI detection, with fairness disparities for EFL learners; (3) overreliance and cognitive offloading, with preliminary evidence of cognitive atrophy; and (4) malleable trustworthiness perceptions, protected by domain knowledge and metacognitive accuracy. The findings reveal a paradox of fluent unreliability: AI-generated errors are often indistinguishable from accurate content, and this risk is inversely distributed with student competence. The themes are integrated into an Epistemic Reliability Framework,  that supports epistemic scaffolding, assessment redesign, and AI literacy targeting metacognition.

İsmail Kaşarcı · 0 citations
Review Open access Sep 2026

Performance and Hallucination Analysis of Large Language Models on European Anesthesiology Examinations: Cross-Sectional Comparative Study

Abstract Background Large language models (LLMs) have shown promising performance on medical examinations across specialties. However, comparative evaluations of current-generation LLMs across multiple European anesthesiology examinations, alongside structured assessment of hallucinations vs question-related confusion, remain lacking. Objective This study aimed to compare the performance of 4 state-of-the-art LLMs on anesthesiology and intensive medicine examination questions and assess their hallucination rates. Methods This computational comparative study analyzed 437 multiple-choice questions (1748 queries) from 3 sources: nurse anesthetist school examinations (infirmier anesthésiste diplômé d’État [registered nurse anesthetist]; n=100, 22.9%), European Diploma in Anaesthesiology and Intensive Care (EDAIC; n=219, 50.1%), and EDAIC On-Line Assessment (n=118, 27.0%). Each question was submitted to 4 LLMs (Claude Sonnet 4.5, Gemini 2.5 Pro, GPT-5, and Grok 4) using standardized prompts via default web interface settings. Responses were evaluated through structured consensus review by 2 examiners for accuracy, hallucinations, and question-related confusion. Statistical analysis included Friedman and Wilcoxon signed-rank tests with Holm-Bonferroni correction, the Cochran Q test, and generalized estimating equations. Results Average success rates ranged from 86% (SD 18%) to 94% (SD 10%) across LLMs and examination types, exceeding the EDAIC part I passing threshold, representing substantial improvement over previously reported GPT-3.5 performance. For the EDAIC, overall intermodel differences were significant (Friedman χ23=13.9; P=.003; W=0.02), with Gemini outperforming GPT-5 as the only pairwise difference. Hallucination rates ranged from 11% (11/100) to 20.1% (44/219) without significant intermodel differences. All models exceeded the EDAIC passing threshold. Conclusions Current-generation LLMs demonstrated consistently high performance across multiple European anesthesiology examinations but continue to produce clinically relevant hallucinations, supporting their role as supervised educational tools rather than autonomous learning resources. These findings underscore the need for structured integration frameworks and systematic verification when deploying LLMs as learning tools in medical education.

Ștefan Andrei, Thibault Giet, Alexis Belouard et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.