Skip to content
Open access

Prompting for Independent Learning: An Evaluation of Tutoring Behaviors in GenAI

Jul 2026 · Future Internet · 0 citations · 28 references

TL;DR

A simulation-based textual analysis of prompt design evaluates a frontier large language model as a tutor across 60 scripted sessions on a single topic and point to dynamic, dialogue-aware prompting alongside explicit SRL scaffolding.

Abstract

Generative AI tutors have become a common tool for independent learning, yet their capacity to support self-regulated learning (SRL) is poorly understood. This simulation-based textual analysis of prompt design evaluates a frontier large language model (Claude Sonnet 4.6) as a tutor across 60 scripted sessions on a single topic (density), crossing three levels of SRL-informed system prompting (Minimal, Moderate, Extensive) with four learner-behavior variants (Standard, Misconception, Disengagement, Overconfidence). Tutoring transcripts were scored on a 14-dimension framework spanning SRL phases, SRL developmental stages, self-determination theory principles, and Merrill’s First Principles of Instruction, applied via an LLM judge. Adding SRL context to the system prompt raised total tutoring scores, but only at the Extensive SRL support level. Minimal and Moderate prompting produced the same performance, near 36 on a 70-point scale, and Extensive prompting raised it to 40, a statistically significant effect (partial η2 = 0.24). The learner’s behavior in the session had a larger effect than the prompt did (partial η2 = 0.37), with disengaged learners scoring lowest. The threshold pattern held under an independent judge from a different developer than the tutor model. The findings support a method for evaluating GenAI tutors empirically and point to dynamic, dialogue-aware prompting alongside explicit SRL scaffolding.

Read PDF

Similar papers

Book Open access Aug 2026

Steering AI Tutors Through System Prompts: A Crossover Study on Self-Regulated Learning and Cognitive Engagement Scaffolds in CS1

Background. Large language models are increasingly deployed as tutors in introductory programming courses, yet evidence that they actually improve learning remains thin, and their tendency to shortcut productive struggle raises concerns about pedagogical harm. Self-regulated learning (SRL) and cognitive engagement (CE) frameworks offer a principled way to address this, but whether embedding them in system prompts actually changes how students learn is an open question. Objectives. We investigated whether AI tutors guided by SRL and CE frameworks affect conceptual understanding, perceived usability, cognitive load, student engagement, and other outcomes when compared to a pedagogically constrained baseline tutor in CS1. Methods. We conducted a preregistered, three-armed crossover study over six weeks of authentic coursework, comparing a baseline AI tutor against two SRL and CE-guided tutors designed to model Zimmerman’s cyclical model, which scaffolds planning, monitoring, and reflection, and Chi’s ICAP framework, which promotes progressively deeper forms of cognitive engagement. We assessed outcomes using post-exercise surveys, conceptual multiple-choice questions, in-platform ratings, and coded responses from the interaction logs, and analyzed the quantitative measures with mixed-effects models. Findings. On the four preregistered confirmatory measures, we found no statistically significant differences between conditions. Non-confirmatory analysis showed that students spent significantly more time on task, wrote longer messages, and produced more constructive contributions when interacting with SRL and CE tutors. Furthermore, the relationship between cognitive load and quiz performance differed significantly by agent type. Implications. Our results suggest that the pedagogical behavior of AI tutors may not be easily steered through system prompts alone: embedding established SRL and CE frameworks did not produce detectable improvements on any preregistered outcome in a large, ecologically valid deployment. Rather than prescribing a single tutoring strategy, future designs may benefit from giving students greater agency over the kind of help they receive, allowing them to choose between scaffolded and more direct support based on their own needs.

Maximilian Georg Barth, Sverrir Thorgeirsson, K. Etemadi et al. · 0 citations
Open access Aug 2026

A GenAI-Based Adaptive Tutoring ana Intelligent Assessment Framework for Personalized Learning

EduMind is introduced, a unified tutoring and assessment platform designed around a dual-track evaluation model that demonstrates how assessment and tutoring can be unified into a seamless workflow, and remained operationally stable throughout all testing phases.

Dhyan Gowda, M. Aruna, P. Prasad et al. · 0 citations
Preprint Aug 2026

LLM Pedagogical Behavior in AI Tutoring Interactions

Students increasingly use LLMs as tutors for coursework and problem solving. Little is known about the level of assistance LLMs provide when students use them as tutors in authentic learning interactions. This matters because tutoring responses can differ substantially in how directly they help students complete a task. We operationalize this dimension as scaffolding level and develop a five-level scale, validated against human annotations, that characterizes responses according to the degree of direct assistance they provide. We apply the scale to 14,637 LLM responses from 203 students in a university AI course. Responses are overwhelmingly concentrated at high levels of assistance, with more than 95% classified as either Explaining or Solving. Scaffolding level is systematically associated with students'subsequent conversational behavior, but provides little additional predictive information about performance on three subsequent exams beyond prior achievement and dialogue behavior. These findings provide an empirical baseline for LLM assistance in tutoring interactions and a measurement framework for evaluating how alternative tutoring designs change that assistance.

Suhyeon Lee, Juneha Baek, Jaehyeong Park et al. · 0 citations
Open access Jul 2026

Toward Standardized Interactivity: An AI-Enabled Adaptive Speaking Task for Computer-Delivered Large-Scale Assessment

Despite task interactivity eliciting different aspects of language use, more interactive, adaptive speaking tasks have been challenging to implement in computer-delivered large-scale high-stakes contexts. Recent advances in large language models (LLMs) offer a potential means to address this tension through spoken dialogue systems (SDSs) but fall short of full adaptivity. To enhance adaptivity in an LLM-enhanced SDS-delivered interview task, this article proposes two solutions: (1) LLM-driven response evaluation using question-specific task completion rubrics for adaptive follow-up question selection and (2) item response theory (IRT) scoring accounting for prompt variability. We examined the extent to which these solutions functioned to simulate adaptivity, from 5,909 test takers’ responses to an adaptive speaking task with an avatar and a non-adaptive monologic task. Automated task completion evaluations were comparable to expert ratings, though follow-up question evaluation showed low agreement both among human raters and between raters and the system. IRT scoring yielded higher test–retest reliability, and the adaptive task elicited more reciprocal language use than the monologic task. Unlike earlier SDSs that prioritized consistency over adaptivity, the real-time meaning-level evaluation with IRT modeling strikes a balance between standardization and adaptivity. These findings support the viability of adaptive speaking tasks in computer-based standardized assessments.

Y. Park, Yigal Attali, Xiaowan Zhang et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.