Skip to content

Author

D. Henderson

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#large language models Open access Aug 2026

How Do We Know They Work? A Systematic Review of Evaluation Practices and Methodological Rigor in LLM-Based Agentic Educational Systems

Background. Large language model (LLM) based agentic educational systems, which plan, use tools, reflect, or coordinate as multiple agents, are proliferating, but whether they are evaluated rigorously enough to support claims about learning is unclear.Objective. To systematically characterize the evaluation practices and methodological rigor of LLM-based agentic educational systems, and to derive a reporting standard for the field.Methods. Following PRISMA 2020, I searched eight sources (Scopus, Web of Science, ACM, IEEE, ERIC, ACL Anthology, arXiv, EdArXiv) for empirical studies of agentic educational systems published from January 2023 to June 2026. A single reviewer screened 3,710 deduplicated records, assisted by a validated advisory LLM triage (blind human-AI agreement Cohen’s κ = 0.65–0.74) that recommended but never determined decisions. Of 477 included studies, the 132 with openly retrievable full text were coded against a 41-field rigor-and-reporting codebook; 129 agentic studies formed the analytic set.Results. Evaluation was predominantly output-oriented: 56.6% used no human learners and only 23.3% measured an actual learning outcome. Rigor and reproducibility were limited: 1.6% used fully validated instruments, 2.3% controlled for novelty effects, 17.8% were fully reproducible, and 50.4% released prompts. Appraisal-critical reporting was sparse (ethics/IRB unreported in 68.2%). Studies disclosed most items (mean 16.2/18) but met far fewer in practice.Conclusions. The field is closer to a transparency norm than a rigor norm. I propose the Agentic Educational AI Evaluation and Reporting (AEER) checklist, a disclosure-focused reporting standard targeting these gaps at low cost. Findings pertain to the open-access subset and are preliminary.

D. Henderson · 0 citations