Jul 2026· IEEE International Conference on Information Reuse and Integration· pp. 86-91· 0 citations· 31 references
Computer Science
TL;DR
The results suggest that learner- and curriculum-aware alignment may matter more for effective tutoring than model category alone, and that such alignment is both measurable and improvable.
Abstract
Large language models (LLMs) are increasingly used as AI tutors, but a correct answer is not always a pedagogically appropriate one. In classroom learning, effective help depends not only on correctness, but also on whether a response matches the learner's current foundation, the course sequence, and the timing of concept introduction. Existing evaluations focus mainly on answer quality, leaving this instructional fit undermeasured. We present the Pedagogical Suitability Index (PSI), a composite metric of six theory-informed sub-scores that evaluates how well LLM-generated tutoring responses align with learner readiness and curricular progression, and we further use PSI as a structured feedback signal for response improvement. We evaluate four LLM tutors (ChatGPT, Gemini, Gemma 4, and Qwen 3) across 240 scenario-based evaluations using paired standard and defective prompts, then apply a PSI-guided regeneration protocol to 62 weak-performing cases. Baseline differences across the four tested models were modest overall (PSI range: 0.557 to 0.638), and open-weight and closed models did not exhibit a clear separation in pedagogical fit. Under the tested prompt perturbations, overall PSI remained largely stable $(\Delta=-0.002)$, though sub-score trade-offs emerged. More importantly, PSIguided feedback substantially improved weak-performing cases: 51 of 62 cases improved (82.3%). Focused manual evaluation of the 62 PSI-selected weak cases provides initial evidence that the identified weaknesses are instructionally meaningful and that many PSI-guided regenerations correspond to human-judged improvement. These results suggest that learner- and curriculumaware alignment may matter more for effective tutoring than model category alone, and that such alignment is both measurable and improvable.
EduClaw-Bench is introduced, a benchmark that places an agent tutor in a continuous 30-day relationship with a simulated learner grounded in knowledge tracing (KT), whose knowledge-concept mastery, from a KT model trained on real-student data, drives its answers and is probed for learning gain across 55 scenarios.
Unggi Lee, Sookbun Lee, Yeil Jeong et al.· 0 citations
Students increasingly use LLMs as tutors for coursework and problem solving. Little is known about the level of assistance LLMs provide when students use them as tutors in authentic learning interactions. This matters because tutoring responses can differ substantially in how directly they help students complete a task. We operationalize this dimension as scaffolding level and develop a five-level scale, validated against human annotations, that characterizes responses according to the degree of direct assistance they provide. We apply the scale to 14,637 LLM responses from 203 students in a university AI course. Responses are overwhelmingly concentrated at high levels of assistance, with more than 95% classified as either Explaining or Solving. Scaffolding level is systematically associated with students'subsequent conversational behavior, but provides little additional predictive information about performance on three subsequent exams beyond prior achievement and dialogue behavior. These findings provide an empirical baseline for LLM assistance in tutoring interactions and a measurement framework for evaluating how alternative tutoring designs change that assistance.
Suhyeon Lee, Juneha Baek, Jaehyeong Park et al.· 0 citations
EduMind is introduced, a unified tutoring and assessment platform designed around a dual-track evaluation model that demonstrates how assessment and tutoring can be unified into a seamless workflow, and remained operationally stable throughout all testing phases.
Dhyan Gowda, M. Aruna, P. Prasad et al.· International Journal of Sci...· 0 citations
Background. Large language models are increasingly deployed as tutors in introductory programming courses, yet evidence that they actually improve learning remains thin, and their tendency to shortcut productive struggle raises concerns about pedagogical harm. Self-regulated learning (SRL) and cognitive engagement (CE) frameworks offer a principled way to address this, but whether embedding them in system prompts actually changes how students learn is an open question. Objectives. We investigated whether AI tutors guided by SRL and CE frameworks affect conceptual understanding, perceived usability, cognitive load, student engagement, and other outcomes when compared to a pedagogically constrained baseline tutor in CS1. Methods. We conducted a preregistered, three-armed crossover study over six weeks of authentic coursework, comparing a baseline AI tutor against two SRL and CE-guided tutors designed to model Zimmerman’s cyclical model, which scaffolds planning, monitoring, and reflection, and Chi’s ICAP framework, which promotes progressively deeper forms of cognitive engagement. We assessed outcomes using post-exercise surveys, conceptual multiple-choice questions, in-platform ratings, and coded responses from the interaction logs, and analyzed the quantitative measures with mixed-effects models. Findings. On the four preregistered confirmatory measures, we found no statistically significant differences between conditions. Non-confirmatory analysis showed that students spent significantly more time on task, wrote longer messages, and produced more constructive contributions when interacting with SRL and CE tutors. Furthermore, the relationship between cognitive load and quiz performance differed significantly by agent type. Implications. Our results suggest that the pedagogical behavior of AI tutors may not be easily steered through system prompts alone: embedding established SRL and CE frameworks did not produce detectable improvements on any preregistered outcome in a large, ecologically valid deployment. Rather than prescribing a single tutoring strategy, future designs may benefit from giving students greater agency over the kind of help they receive, allowing them to choose between scaffolded and more direct support based on their own needs.
Maximilian Georg Barth, Sverrir Thorgeirsson, K. Etemadi et al.· International Computing Educ...· 0 citations
Interest-based learning (IBL) is an educational approach where learners’ interests are used to contextualize learning. IBL can make instruction feel more relevant and lead to improved learning outcomes, but it is difficult for instructors to implement at scale because learner interests are highly varied. Large language models (LLMs) can support IBL through conversational AI tutors that personalize instruction to individual interests. This paper presents a prompt design approach for creating LLM tutors for IBL. We first conducted a literature review to derive pedagogy-grounded strategies for a base tutor prompt, then embedded additional IBL strategies to produce an IBL tutor prompt. We evaluated both prompts via expert review and a human-participants study with undergraduate students. Results show the IBL prompt reliably integrated learner interests, but exhibited shallow reflection, inconsistent knowledge checks, and surface-level analogies when interests were underspecified. We contribute a reusable prompt design pipeline, prompt templates, and evaluation artifacts for designing interest-based AI tutors.
Abhishek Kulkarni, S. Brown, Neha Rani et al.· International Conference on...· 0 citations
A simulation-based textual analysis of prompt design evaluates a frontier large language model as a tutor across 60 scripted sessions on a single topic and point to dynamic, dialogue-aware prompting alongside explicit SRL scaffolding.
Kendall Hartley, Fabiola Sáez-Delgado, Javier Mella-Norambuena· Future Internet· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.