The layer of lesson preparation delegated to AI predicts classroom discourse quality: a four-wave video study of 90 junior secondary mathematics teachers
Abstract
Research on teachers and generative artificial intelligence has largely measured how much they use these tools, leaving open which cognitive operations they hand over. We asked whether the layer of processing delegated during lesson preparation predicts what happens in the classroom. Ninety junior secondary mathematics teachers in 18 public schools in eastern China were followed across four waves of video-recorded lessons over one school year, yielding 897 coded lessons, 4,640 unplanned student contributions, 1,518 preparation sessions, and 556 stimulated-recall interviews. For each lesson, the final plan was compared against the interaction logs segment by segment, producing two ratios that index offloading by provenance: generative-layer offloading, the share of questioning and variation segments originating in model output, and formulation-layer offloading, the share of context and wording segments. Three-level models decomposing within- and between-teacher variation showed that the two layers behaved differently. Within teachers, generative-layer offloading was associated with less higher-order questioning, less substantive uptake of student talk, weaker maintenance of high cognitive demand, and lower recognition and substantive use of unplanned student contributions, with standardized coefficients between 0.08 and 0.18. Formulation-layer offloading was unrelated to four of those indicators, positively associated with the length and variety of designed variation sequences, and positively associated with recognition of unplanned contributions, the single indicator on which the two layers cross in opposite directions. Total preparation time with AI predicted discourse outcomes on its own and lost that value for the lesson-level indicators once layer composition was entered. An embedded micro-randomized comparison, in which teachers drafted their own questioning and variation sequence before consulting the model, lowered generative-layer offloading and raised substantive uptake. Professional attention to student thinking, coded from stimulated-recall interviews conducted after the lesson by a separate team, followed the same within-teacher pattern and carried the association with the substantive use of unplanned contributions. However, the retrospective measurement leaves its direction open. Teachers’ post-preparation confidence rose with generative-layer offloading by 0.23 standard deviations per within-teacher standard deviation. It carried no positive information about the lesson that followed, a confidence-performance dissociation. What predicted classroom discourse was which layer was delegated, not how much.