Skip to content
Review Open access

Effects of AI-generated feedback on l2 writing: a PRISMA 2020 systematic review of performance, revision, and engagement outcomes

Sep 2026 · JRTI (Jurnal Riset Tindakan Indonesia) · Vol 11, pp. 998-1013 · 0 citations · 69 references

TL;DR

The evidence indicates that GenAI feedback most consistently supports local accuracy, lexical development, revision activity, and learner engagement, whereas effects on syntactic complexity, higher-order composing, and long-term retention are mixed or conditional.

Abstract

Generative artificial intelligence (GenAI) tools can provide immediate feedback during second-language (L2) writing, but evidence on their effects across writing outcomes remains fragmented. This PRISMA 2020 systematic review synthesized peer-reviewed empirical studies published from 2023 to 2025 that examined GenAI-generated feedback for English writing in EFL/ESL contexts. Searches of Scopus, PubMed, and ScienceDirect identified 1,106 records; after 12 duplicates were removed, 1,094 records were screened, 71 full texts were assessed, and 20 studies met the eligibility criteria. Eligible studies were experimental, quasi-experimental, mixed-methods, qualitative, or design-based investigations reporting writing performance, revision, feedback engagement/literacy, or AI-feedback accuracy. Because study designs, writing tasks, outcome measures, and reported statistics were heterogeneous, statistical meta-analysis and pooled effect-size estimation were not conducted; findings were synthesized thematically and by effect direction. The evidence indicates that GenAI feedback most consistently supports local accuracy, lexical development, revision activity, and learner engagement, whereas effects on syntactic complexity, higher-order composing, and long-term retention are mixed or conditional. Comparative evidence generally supports a complementary model in which AI provides rapid, high-volume local feedback and teachers address global structure, disciplinary expectations, contextual appropriateness, and affective support. Key limitations of the evidence base include short intervention periods, small samples, heterogeneous measurement, limited longitudinal evidence, and uneven geographic representation. Future research should preregister screening procedures, use independent double screening with agreement statistics, report extractable effect sizes and delayed post-tests, and examine equity, transfer, and hybrid AI-human feedback models.

Read PDF

Similar papers

Open access Sep 2026

Do AI Writing Tools Improve L2 Writing Performance? A Meta-Analysis of Experimental Evidence, Moderators, and Methodological Quality

Artificial intelligence (AI) writing tools—including large language model (LLM)-based systems, automated writing evaluation (AWE) platforms, and AI-powered writing assistants—have become an increasingly prominent feature of second language (L2) writing pedagogy. Yet whether these tools reliably improve L2 writing perfo...

Jin-Ming Du · 0 citations
Review Open access Sep 2026

Conversational AI for mental health: a systematic review of effectiveness and design

This systematic review synthesised evidence on the applications, effectiveness, engagement, acceptability, and responsible design of large language model (LLM)- and AI-based conversational agents for mental health and psychological well-being. Following the PRISMA 2020 framework, Scopus, PubMed, and ScienceDirect were...

Anisa Sholiha Mia, Rudyanto Rudyanto · 0 citations
#small language model Review Open access Sep 2026

Conditional validity in LLM-mediated L2 assessment: an argument-based systematic review and meta-analysis

Introduction Large language models (LLMs) are increasingly used for scoring and feedback in second-language (L2) assessment, yet the validity of the resulting interpretations remains contested. This review evaluated when LLM-mediated assessment is psychometrically and educationally defensible using an argument-based va...

L. Alghamdi, T. Alghizzi · 0 citations
Review Open access Aug 2026

A Systematic Review of Empirical Research about Beyond Automated Correction by Human–AI Feedback Partnership for Developing L2 Writing

A systematic review of empirical studies published from 2023 to 2026 demonstrates that generative AI facilitates writing quality improvement, revision processes, learner autonomy and increased writing access by providing on-time, tailored and interactive support.

Muhammad Imran, Z. Ali, Mohammad Musab Bin Azmat Ali · 0 citations
Review Open access Sep 2026

Does AI Help EFL Learners? A Scoping Review of Effectiveness, Engagement, and Ethics Across 131 Scopus-Indexed Studies

- This scoping review maps 131 Scopus-indexed publications (2022 – 2026) examining AI-mediated English language teaching (ELT). Following PRISMA 2020 guidelines (Page et al., 2021), we searched Scopus and analyzed retrieved records through bibliometric analysis, thematic synthesis, and keyword co-occurrence mapping. Pu...

S. Bharathikumar, B. K · 0 citations
#artificial intelligence Review Open access Sep 2026

Artificial intelligence in mathematics education: A PRISMA-based systematic literature review (2021-2025)

Overall, AI shows potential to support mathematics teaching and learning, but stronger longitudinal and comparative evidence is needed to establish effectiveness, equity, and sustainable implementation.

F. Omirzakova, Sarsenkul Tileubay · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.