Artificial Intelligence in Language Education in the Generative-AI Era: A Systematic Integrative Review of Pedagogical Evidence, Methodological Quality, and Responsible Adoption
Abstract
Artificial intelligence has grown quickly in language teaching, particularly since the introduction of generative AI and large language models. Even so, the existing evidence is dispersed across various technologies, teaching applications, learner groups, and research approaches. This systematic integrative review examined 40 primary empirical or technical studies. It also used 15 systematic reviews, meta-analyses, and bibliometric syntheses to provide contextual evidence. The selected publications were published between 2017 and July 2026, while the majority of the primary evidence appeared after 2022. The review employed descriptive mapping, thematic synthesis, design-sensitive methodological assessment, and framework development. The primary papers were evaluated based on AI technology, educational application, language skill, learner group, research design, outcome type, multilingual representation, methodological quality, and ethical or governance concerns. Writing assistance, automated feedback, speaking practice, learner perceptions, teacher support, and AI-assisted pedagogy were the most extensively researched topics. In contrast, reading, listening, younger learners, low-resource languages, delayed learning outcomes, and validation across contexts received less attention. The evidence points to potential short-term improvements for writing revision, speaking practice, motivation, engagement, and teacher support. However, confidence in these findings is limited because many studies used small sample sizes, brief interventions, single-institution settings, self-reported measures, incomplete model and prompt reporting, and minimal external validation. Generative systems increased interactivity, adaptability, and task coverage but also introduced output variability, hallucination, authorship ambiguity, privacy risk, and reproducibility problems. Based on the integrated findings, the review proposes a responsible-adoption framework connecting AI-system characteristics, pedagogical design, human–AI interaction, contextual conditions, educational outcomes, and institutional governance. The findings indicate that educational value depends less on technological novelty than on task alignment, teacher mediation, learner agency, methodological rigour, multilingual inclusion, and accountable human oversight.