Claims that generative artificial intelligence (GenAI) improves language learning depend on the counterfactual, the assessment condition, and the time horizon. We conducted a convergent segregated mixed-methods evidence audit of GenAI-supported English-language learning and introduced an inference-ceiling framework that matches each claim to its minimum design requirements. The verified database consolidated 676 source-study rows into 530 canonical records; 205 studies met direct scope, including 29 controlled studies (N = 3,231), 78 variance-complete controlled effects, 545 direct-scope mechanism findings from 96 studies, and 51 direct-scope mixed-methods bridges from 46 studies. Random-effects models used restricted maximum likelihood, Hartung-Knapp confidence intervals, and prediction intervals. Objective language performance favored GenAI-supported packages over usual practice (k = 10, g = 1.23, 95% CI [0.44, 2.03], prediction interval [-1.00, 3.47]) but was smaller and imprecise against established active alternatives (k = 7, g = 0.71, 95% CI [-0.06, 1.48], prediction interval [-1.41, 2.83]). A post hoc active-minus-usual coefficient was -0.51 (95% CI [-1.54, 0.53], p = .316), so descriptive attenuation was not statistically distinguishable from zero. After full-text verification, assessment access remained unreported in the source for 74 of 78 controlled effects; only one explicitly used tool-withdrawn assessment, and no variance-complete effect identified GenAI's incremental contribution within a common instructional base. Study-clustered mechanism sensitivity showed that negative or boundary evidence appeared in 76.2% of studies contributing to the offloading/integrity family. Current evidence supports context-dependent package benefits, not a stable estimate of GenAI-specific or independently retained learning.
Wen Hou, Nan Li, Akbar Bahari· Zenodo (CERN European Organi...· 0 citations
Claims that generative artificial intelligence (GenAI) improves language learning depend on the counterfactual, the assessment condition, and the time horizon. We conducted a convergent segregated mixed-methods evidence audit of GenAI-supported English-language learning and introduced an inference-ceiling framework that matches each claim to its minimum design requirements. The verified database consolidated 676 source-study rows into 530 canonical records; 205 studies met direct scope, including 29 controlled studies (N = 3,231), 78 variance-complete controlled effects, 545 direct-scope mechanism findings from 96 studies, and 51 direct-scope mixed-methods bridges from 46 studies. Random-effects models used restricted maximum likelihood, Hartung-Knapp confidence intervals, and prediction intervals. Objective language performance favored GenAI-supported packages over usual practice (k = 10, g = 1.23, 95% CI [0.44, 2.03], prediction interval [-1.00, 3.47]) but was smaller and imprecise against established active alternatives (k = 7, g = 0.71, 95% CI [-0.06, 1.48], prediction interval [-1.41, 2.83]). A post hoc active-minus-usual coefficient was -0.51 (95% CI [-1.54, 0.53], p = .316), so descriptive attenuation was not statistically distinguishable from zero. After full-text verification, assessment access remained unreported in the source for 74 of 78 controlled effects; only one explicitly used tool-withdrawn assessment, and no variance-complete effect identified GenAI's incremental contribution within a common instructional base. Study-clustered mechanism sensitivity showed that negative or boundary evidence appeared in 76.2% of studies contributing to the offloading/integrity family. Current evidence supports context-dependent package benefits, not a stable estimate of GenAI-specific or independently retained learning.
Wen Hou, Nan Li, Akbar Bahari· Zenodo (CERN European Organi...· 0 citations
Artificial-intelligence tools now deliver feedback to English language learners through automated writing evaluation, speech-recognition systems, intelligent tutoring, and generative chatbots, but evidence on their effects remains uneven. This convergent mixed-methods systematic review synthesised 195 studies of AI-mediated feedback in English as a foreign or second language published between 2023 and 2026, combining random-effects meta-analysis of comparative performance effects with thematic synthesis of learner-uptake evidence. Seventeen studies contributed an eligible comparative effect on objective language performance. The pooled estimate was large and highly heterogeneous, Hedges’ g = 1.19, 95% CI [0.57, 1.81], τ² = 1.42, I² = 93.2%, with a 95% prediction interval from −1.34 to 3.72. Restricting the pool to effects below g = 2 produced a more stable model, g = 0.62 [0.40, 0.85], k = 13, I² = 52.6%. Domain estimates were unstable: writing, g = 1.36 [0.26, 2.46], k = 9, carried I² = 95.7%, while speaking and grammar confidence intervals included zero. Thematic synthesis of study-level summaries identified noticing, uptake and revision, scaffolding and practice, and affective response as recurring proximal mechanisms. All fitted models had positive point estimates; this descriptive pattern does not resolve bias, comparator differences, or uncertainty across settings. The estimate’s magnitude depends heavily on which effects are admitted, and the corpus rarely records whether outcomes were assessed with the AI tool available, so improvement in an assisted product cannot be separated from change in independent performance. Every contributing effect has been traced to a named outcome and contrast in its primary report, and the eight records whose extraction was corrected during revision are logged with their source basis in the supplement.
Meng Zhang, Xiu Xin, Akbar Bahari· Zenodo (CERN European Organi...· 0 citations
Generative artificial intelligence (GenAI) is often treated as a single educational intervention even when studies compare different guidance structures and assess outcomes under different levels of tool availability. This mixed-methods systematic-review draft examined comparison conditions, self-regulatory processes, and independent performance in a source-located corpus of 94 full-text records. The 93 reports yielded 94 study units, 80 provisionally eligible studies, 129 arm configurations, 409 outcome points, 1,338 quantitative results, 412 process-evidence records, and 300 qualitative-evidence records. Among 156 performance-classifiable outcomes, GenAI was absent during 41 assessments, present during 31, arm-specific or mixed during 22, and unreported or unclear during 62. A prespecified pooling gate rejected meta-analysis because 103 metrics, five non-equivalent estimand families, sparse compatible cells, and structural dependence prevented a construct-valid pooled effect. Fourteen unique studies measured independent performance: four were favourable, seven null or mixed, two adverse or configuration-dependent, and one a within-GenAI guidance contrast. Monitoring, strategic prompting, and revision were well represented, whereas metacognitive offloading and illusion of understanding were rarely operationalized. Integrated findings suggest that guidance requiring learners to attempt, evaluate, or revise may preserve independent performance more reliably than unrestricted access, but the evidence matrix is dominated by missing assessment-state information. Every primary outcome should report whether eligible generative functions were available, what functions were available, and how restrictions were enforced.
Zhengbin Dong, Yue Qiu, Akbar Bahari· Zenodo (CERN European Organi...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.