Claims that generative artificial intelligence (GenAI) improves language learning depend on the counterfactual, the assessment condition, and the time horizon. We conducted a convergent segregated mixed-methods evidence audit of GenAI-supported English-language learning and introduced an inference-ceiling framework that matches each claim to its minimum design requirements. The verified database consolidated 676 source-study rows into 530 canonical records; 205 studies met direct scope, including 29 controlled studies (N = 3,231), 78 variance-complete controlled effects, 545 direct-scope mechanism findings from 96 studies, and 51 direct-scope mixed-methods bridges from 46 studies. Random-effects models used restricted maximum likelihood, Hartung-Knapp confidence intervals, and prediction intervals. Objective language performance favored GenAI-supported packages over usual practice (k = 10, g = 1.23, 95% CI [0.44, 2.03], prediction interval [-1.00, 3.47]) but was smaller and imprecise against established active alternatives (k = 7, g = 0.71, 95% CI [-0.06, 1.48], prediction interval [-1.41, 2.83]). A post hoc active-minus-usual coefficient was -0.51 (95% CI [-1.54, 0.53], p = .316), so descriptive attenuation was not statistically distinguishable from zero. After full-text verification, assessment access remained unreported in the source for 74 of 78 controlled effects; only one explicitly used tool-withdrawn assessment, and no variance-complete effect identified GenAI's incremental contribution within a common instructional base. Study-clustered mechanism sensitivity showed that negative or boundary evidence appeared in 76.2% of studies contributing to the offloading/integrity family. Current evidence supports context-dependent package benefits, not a stable estimate of GenAI-specific or independently retained learning.
Wen Hou, Nan Li, Akbar Bahari· Zenodo (CERN European Organi...· 0 citations
Claims that generative artificial intelligence (GenAI) improves language learning depend on the counterfactual, the assessment condition, and the time horizon. We conducted a convergent segregated mixed-methods evidence audit of GenAI-supported English-language learning and introduced an inference-ceiling framework that matches each claim to its minimum design requirements. The verified database consolidated 676 source-study rows into 530 canonical records; 205 studies met direct scope, including 29 controlled studies (N = 3,231), 78 variance-complete controlled effects, 545 direct-scope mechanism findings from 96 studies, and 51 direct-scope mixed-methods bridges from 46 studies. Random-effects models used restricted maximum likelihood, Hartung-Knapp confidence intervals, and prediction intervals. Objective language performance favored GenAI-supported packages over usual practice (k = 10, g = 1.23, 95% CI [0.44, 2.03], prediction interval [-1.00, 3.47]) but was smaller and imprecise against established active alternatives (k = 7, g = 0.71, 95% CI [-0.06, 1.48], prediction interval [-1.41, 2.83]). A post hoc active-minus-usual coefficient was -0.51 (95% CI [-1.54, 0.53], p = .316), so descriptive attenuation was not statistically distinguishable from zero. After full-text verification, assessment access remained unreported in the source for 74 of 78 controlled effects; only one explicitly used tool-withdrawn assessment, and no variance-complete effect identified GenAI's incremental contribution within a common instructional base. Study-clustered mechanism sensitivity showed that negative or boundary evidence appeared in 76.2% of studies contributing to the offloading/integrity family. Current evidence supports context-dependent package benefits, not a stable estimate of GenAI-specific or independently retained learning.
Wen Hou, Nan Li, Akbar Bahari· Zenodo (CERN European Organi...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.