Jul 2026
Beyond Direct Answering: Aligning Educational LLMs as Socratic Guides via Heuristic Reinforcement Learning
HeuristicEdu is presented, a two-phase pipeline that aligns Qwen2.5-7B toward Socratic tutoring via supervised warm-up and Group Relative Policy Optimization (GRPO), and Scaffolding Effectiveness and Conversation Depth are introduced to evaluate outcomes beyond surface fluency.
Xiaokun Wang, Siyu Song, Wentao Liu et al.
· arXiv.org · 0 citations