Learning Analytics of Students’ Interaction with ChatGPT in Programming Education: A Process-Oriented Analysis
Abstract
Generative AI tools are now widely used in undergraduate programming, yet most evidence about how students use them comes from self-report rather than from observed behaviour. This study examined the sequential structure of students’ ChatGPT (GPT- 4o, OpenAI)-supported programming work and the cognitive complexity of their queries. Screen recordings of 363 second-year students completing an individual Python 3.12 data-visualisation assignment were coded into 3985 activity episodes and analysed using descriptive statistics, lag-1 sequential analysis, and cognitive network analysis; because recordings capture actions rather than cognition, the coded categories are treated as behavioural indicators interpreted within, rather than as measurements of, the Self-Regulated Learning framework. Programming activities accounted for 63% of coded actions and ChatGPT interactions for 20%. Behaviour was organised around a troubleshooting cycle, the strongest association being between submitting error messages and reviewing ChatGPT responses (PCM → RF, Yule’s Q = 0.84, a descriptive association measure, rather than a transition probability, whose stability across students was not tested). ChatGPT use was concentrated in activities indicative of monitoring and control and was largely absent from planning and reflection. High-achieving students produced a higher proportion of deep-level submissions (32.2% vs. 19.2%) and a lower proportion of surface-level submissions (23.2% vs. 38.4%) than low-achieving students; because submissions are nested within students, this difference is reported as a property of the observed distribution rather than as an inferential finding. Comparisons computed at the level of the student were tested inferentially and reached significance with small effect sizes; comparisons computed at the level of coded actions or query submissions are reported throughout as observed properties of the aggregate distributions rather than as inferentially established differences. These episode-level findings support instructional scaffolding that structures query formulation and reflection.