HeuristicEdu is presented, a two-phase pipeline that aligns Qwen2.5-7B toward Socratic tutoring via supervised warm-up and Group Relative Policy Optimization (GRPO), and Scaffolding Effectiveness and Conversation Depth are introduced to evaluate outcomes beyond surface fluency.
Abstract
Large language models (LLMs) deployed in educational settings often behave as direct answerers: they disclose target concepts in the opening turn instead of guiding students through progressive inquiry, as Socratic pedagogy prescribes. We present HeuristicEdu, a two-phase pipeline that aligns Qwen2.5-7B toward Socratic tutoring via supervised warm-up and Group Relative Policy Optimization (GRPO). Training uses SocraticEdu, 797 multi-turn Chinese children's science dialogues reconstructed from a live platform, with a heuristic reward over cognitive depth (R_cog), curiosity engagement (R_eng), and directness (R_dir), together with a K_query correction for student-introduced terms. We introduce Scaffolding Effectiveness (SE) and Conversation Depth (CD) to evaluate outcomes beyond surface fluency. On 30 held-out questions, the best GRPO variant improves SE from 30.0% to 63.3% and lowers keyword leakage from 30.0% to 13.3%. Notably, this best variant omits the directness penalty during optimization, suggesting that explicit anti-leakage terms can conflict with gradient-based behavioral alignment. An unaligned Qwen-72B baseline reaches 0% SE and 96.7% leakage, showing that scale alone does not induce Socratic behavior.
Michael, a syllabus-aware AI teaching assistant designed to scaffold reasoning through structured, hint-first dialogue aligned with course progression, rather than providing direct solutions, is introduced, suggesting that curriculum-aligned constraints and hint-first scaffolding can support instructional integration without displacing pedagogical goals.
Or Peretz, Roei Zerahia· International Journal of Inf...· 0 citations
This research addresses the common challenge of a lack of context in Question and Answer (QA) datasets in digital education, which limits the reasoning potential of Large Language Models (LLMs). To address this, we optimize an automated retrieval-based dataset generation system that systematically enriches QA pairs with relevant pedagogical context from authoritative digital textbooks. This study conducts a comparative analysis of two major text chunking strategies: sentence chunking and recursive chunking. Although these pipelines are designed for general education applications, they are evaluated here through a case study of Indonesian elementary education materials. To ensure the highest reliability, the workflow performance is measured against a ground truth dataset of 978 entries, manually curated and validated by education experts to ensure pedagogical accuracy, and 781 entries from other subjects. Quantitative evaluation using BERTScore shows that recursive chunking achieves a superior F1 score of 0.748 compared to 0.737 for sentence chunking, with peak performance observed on upper elementary school materials (Grades 5 and 6). These findings were corroborated by the final verification phase through User Acceptance Testing (UAT) with an elementary school educator, where recursive chunking achieved a 'Relevant' score of 22 compared to 17 for sentence chunking. A key contribution of this study is the development and validation of a standardized, automated workflow by experts that effectively overcomes the barriers of manual dataset construction for domain-specific tasks, providing a semantically robust foundation for context-aware educational AI.
V. C. Mawardi, Ayu Purwarianti, B. Trilaksono et al.· International Conference on...· 0 citations
Large language models (LLMs) often exhibit sycophancy, optimizing for agreement over productive challenge, which severely limits their utility in domains like professional skills training, where growth requires pushback. We introduce, ConvoDojo, a novel conversational AI platform for practicing difficult workplace conversations, engineered not merely as a commercial training application but also as a flexible, instrumented research platform for evaluating conversational AI strategies. ConvoDojo repurposes LLMs as structured sparring partners to support skill development in difficult workplace conversations (e.g., performance feedback, conflict resolution), addressing the reported managerial tendency to avoid them. This paper showcases the platform and presents an evaluation of how key conversational user interface (CUI) design elements, namely, the addition of structured feedback and upfront instructional scaffolding, impact managers’ learning. Results show that ConvoDojo is highly engaging and promotes user reflection. We demonstrate how theory-informed dialogue and adaptive pushback can transform an LLM into an effective, measurable tool for complex communication skills development.
Everlyne Kimani Cross, Luiza A Santos, Laurent Denoue et al.· International Conference on...· 0 citations
Socratic AI, a VS Code-integrated tutor that addresses this through pedagogically-grounded Socratic dialogue constrained to withhold direct solutions is presented, and evidence that stateful tracking enables adaptive Socratic dialogue that scaffolds productive struggle rather than short-circuiting learning is contributed.
Ayush Thonge, Aalok Thakkar· Annual Conference on Innovat...· 0 citations
EduMind is introduced, a unified tutoring and assessment platform designed around a dual-track evaluation model that demonstrates how assessment and tutoring can be unified into a seamless workflow, and remained operationally stable throughout all testing phases.
Dhyan Gowda, M. Aruna, P. Prasad et al.· International Journal of Sci...· 0 citations
In higher education, debates about Generative Artificial Intelligence (GenAI) often polarize around academic integrity risks and efficiency gains. A growing body of post-2023 work has begun to move beyond this dichotomy, proposing constrained tutoring systems, Socratic dialogue agents, and adaptive pedagogical scaffolds built on large language models (LLMs). However, these efforts typically target a single instructional function (e.g., Socratic questioning, problem-by-problem scaffolding, or guardrailed answer generation) and treat the cognitive demands of learning as undifferentiated. An instructional architecture that explicitly aligns distinct LLM configurations with the qualitatively different cognitive operations required across learning phases is not available in the literature yet. We address this gap by presenting a simulator-based framework that decomposes instruction into three functionally distinct, sequentially gated simulators: (i) structured comprehension with explicit depth regulation, (ii) schema-based application and analysis under progressively increasing demands, and (iii) evaluation and creation under instructor-defined epistemic uncertainty. Each simulator is grounded in a specific cognitive theory (Zone of Proximal Development, dual-process accounts of cognition, and epistemic cognition, respectively) and is operationalized through explicit constraints, transition criteria, and non-normative diagnostic rubrics. The framework conceptualizes LLMs as constrained instructional simulators whose pedagogical value derives from how they are configured, bounded, and sequenced by the instructor. The primary contribution is therefore not a new tutoring paradigm but a theory-aligned architecture for sequencing multiple, functionally distinct LLM configurations within a single instructional design. Rather than asserting a solution to Bloom's 2-Sigma Problem, the framework demonstrates how LLMs can scale the specific instructional mechanisms associated with individualized tutoring while preserving disciplinary standards and pedagogical authority. The framework is conceptual and design-oriented; empirical validation is identified as a necessary next step.
C. Ugrinowitsch, C. Libardi· Frontiers in Education· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.