Skip to content
Review Open access

Retrieval-augmented generation for pedagogically aware educational AI: an expert-rated comparison of a prompt-only LLM tutor and an integrated, learner-state-aware RAG tutor

Aug 2026 · Frontiers in Education · 0 citations · 23 references

Abstract

Large language models can produce fluent tutoring dialogue, but their educational use remains limited by weak grounding, uneven pedagogical control, and the risk of unsupported content. This study compared two controlled algebra tutoring workflows built on the same foundation model, Gemini 2.5 Flash. The first condition was a prompt-only tutor that used the immediate task and scripted follow-up only. The second was an integrated pedagogical RAG tutor that combined a curated algebra corpus, learner-state variables, and a pedagogical response policy. In total, 24 two-turn algebra episodes were evaluated. No students were recruited, no classroom intervention was conducted, and no learning outcomes were measured. Eight expert reviewers scored anonymized response pairs for solution accuracy, conceptual support, instructional quality, instructional helpfulness, and unsupported content. The pedagogical RAG condition received higher task-level expert ratings for conceptual support (M = 4.69 vs. 4.31), instructional quality (M = 4.61 vs. 4.29), and instructional helpfulness (M = 4.65 vs. 4.39). It also achieved a higher pedagogical-quality composite score (M = 4.65 vs. 4.33; Wilcoxon p  = .00024; rank-biserial r = .86). Solution accuracy was high in both conditions and showed only a small descriptive difference (M = 4.84 vs. 4.74; p  = .0769). However, unsupported content flags were more frequent in the RAG condition, occurring in 13 of 192 rating cells compared with one of 192 rating cells in the baseline condition. Fisher's exact test yielded p  = .0015. These findings do not establish learning gains, classroom efficacy, or the separate causal value of retrieval alone. They indicate that an integrated learner-state-aware RAG architecture can receive stronger expert ratings for pedagogical text quality on a fixed task set while also introducing trace leakage and evidence-boundary risks that must be controlled before deployment

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.