What predicts second-language vocabulary retention? A common-metric comparison of behavioural learning traces and self-reported strategy use
Abstract
Research on second-language vocabulary retention has typically drawn on two largely distinct measurement traditions: learner self-reports of strategy use collected using questionnaires, and behavioural records of practice collected by digital learning platforms. Findings from the two traditions are often cited interchangeably, yet no study has compared their predictive value for retention using a common effect size metric at the level of the individual learner. We report three linked analyses framed as estimation rather than significance contrast. Study 1 modelled 1,930,889 practice traces from 17,230 learners of six languages in a public Duolingo dataset under leakage-strict, user-level cross-validation. Study 1b aggregated the same data to the person level in a prospective design, predicting week-2 recall from week-1 behaviour for 3,957 learners. Study 2 measured 109 English as a Foreign Language (EFL) learners with a 45-item strategy questionnaire and a two-wave vocabulary test, analysed under four pre-defined specifications with permutation-based inference. Recorded practice history and lag predicted recall (gradient-boosting area under the ROC curve = 0.607 ± 0.009; Spearman ρ = 0.138 ± 0.013), with near-perfect calibration (expected calibration error = 0.002) and robustness to the exclusion of immediate-retest traces. At the person level, week-1 behaviour predicted week-2 recall [practice-only features, r = 0.195, 95% CI (0.165, 0.225); with week-1 accuracy added, r = 0.469 (0.445, 0.493)]. Self-reported strategy use did not predict measured retention [primary r = −0.108 (−0.290, 0.082)]: none of the 36 subscale-level and none of the 45 item-level associations survived false-discovery-rate correction, and cross-validated classification did not exceed chance (permutation p ≥ 0.43). The upper confidence bound of the self-report estimate lies below the lower bounds of both person-level behavioural estimates, and this ordering survives correction for measurement unreliability. We discuss implications for measurement choices in instructed second-language acquisition research.