Reliability Risks of Generative AI in Education: A Systematic Review of Hallucinations, Misinformation, Overreliance, and Assessment Validity
Generative AI (GenAI) has introduced reliability risks in education, including hallucinations, fabricated citations, factual inaccuracies, overreliance, and threats to assessment validity. Empirical evidence on these risks has not been systematically synthesized. This review synthesizes that evidence and its consequences for learning and assessment across four research questions. Following PRISMA 2020, Web of Science and Scopus were searched. After duplicate removal, 155 records were screened by two independent doctoral researchers (Cohen's κ = .847). Thirty-five empirical studies were included; 32 (91.4%) were rated high quality using the Mixed Methods Appraisal Tool (MMAT). Thematic synthesis identified four themes: (1) hallucinations and factual inaccuracies taken up as misinformation, moderated by domain expertise; (2) reliability limitations in automated essay scoring, classroom observation, and AI detection, with fairness disparities for EFL learners; (3) overreliance and cognitive offloading, with preliminary evidence of cognitive atrophy; and (4) malleable trustworthiness perceptions, protected by domain knowledge and metacognitive accuracy. The findings reveal a paradox of fluent unreliability: AI-generated errors are often indistinguishable from accurate content, and this risk is inversely distributed with student competence. The themes are integrated into an Epistemic Reliability Framework, that supports epistemic scaffolding, assessment redesign, and AI literacy targeting metacognition.