The Teaching Monster Challenge is introduced, the first instructional video generation benchmark to treat the learner persona as an explicit evaluation criterion and release the benchmark, rubric, and human judgments as a testbed for both teaching systems and automatic judges.
Abstract
AI agents can now solve problems, answer like subject experts, and generate long-form multimodal content. However, whether they can adapt a lesson to fit a specified learner, which education calls Pedagogical Content Knowledge (PCK), has not been benchmarked. To measure it, we introduce the Teaching Monster Challenge, the first instructional video generation benchmark to treat the learner persona as an explicit evaluation criterion. Each system is given a topic and a learner persona and must generate a complete instructional video. Every video is screened by an LLM-judge, ranked by crowd pairwise voting, and finalized by an expert panel. The first edition shows that today's systems handle the content well but are far weaker at presenting it and adapting it to the learner. The same process exposes a limit of automatic judging. The LLM-judge separates a clear low-performing tail but ranks the strongest systems poorly. The strongest systems receive nearly identical scores from the judge, so its ranking of them does not match human preference. Progress therefore requires not only better teaching systems but also better automatic judges, and we release the benchmark, rubric, and human judgments as a testbed for both.
EduClaw-Bench is introduced, a benchmark that places an agent tutor in a continuous 30-day relationship with a simulated learner grounded in knowledge tracing (KT), whose knowledge-concept mastery, from a KT model trained on real-student data, drives its answers and is probed for learning gain across 55 scenarios.
Unggi Lee, Sookbun Lee, Yeil Jeong et al.· 0 citations
LLMs have achieved remarkable capabilities in generating language teaching slides. However, a critical mismatch persists between visual polish and actual instructional effectiveness. To address this gap, we introduce SLATE (Slide-based Learning Assessment for Teaching Effectiveness), the first benchmark that evaluates AI-generated language teaching slides through instructional effectiveness and learner knowledge acquisition. SLATE transforms linguistics olympiad puzzles from low-resource languages with negligible web presence into 90 standardized instructional units comprising 1,133 assessable items, paired with a structured course outline and matched near- and far-transfer test sets. This pretest-posttest design eliminates pretrained knowledge leakage, ensuring gains reflect learning rather than prior recall. Using VLMs as scalable learner proxies and directionally supported by a three-system human pilot, our results show that content validity exhibits a weak association with learning gain, while pedagogical design exhibits a robust positive association. Moreover, most systems show a significant gap between near- and far-transfer accuracy, and even frontier models can produce negative learning gains. SLATE reveals a dissociation between artifact quality and instructional effectiveness, calling for a paradigm shift in how generative teaching systems are built, evaluated, and deployed.
Jing-Zhuo Wu, Jia-Jun Zhang, Liu Yi et al.· 0 citations
The results suggest that learner- and curriculum-aware alignment may matter more for effective tutoring than model category alone, and that such alignment is both measurable and improvable.
Benjamin Barlog, Hudson Craig, Ze-Dong Peng· IEEE International Conferenc...· 0 citations
Autonomous artificial intelligence (AI) agents can now log into a learning management system, read course materials, and complete unproctored, asynchronous assessed work end-to-end with no student involvement. We document that capability and trace its consequences for assessment validity. Three demonstrations on a live undergraduate course supply the evidence: two quiz completions, one in approximately 12 minutes, one in under 5, and a third in which the agent fabricated credible personal reflection for a discussion board. The wider public record includes at least 15 documented agent runs across three platforms and seven tools. We apply Kane’s argument-based validity framework: agent completion removes the attribution on which every inference in Kane’s chain depends. Everything downstream, from course grades to the evidence chains behind program review and accreditation, rests on support that is no longer there. The failure concerns validity rather than integrity: an institution can punish misconduct and still lack grounds for the scores it reports. Collective accreditor guidance addresses institutional uses of AI in evaluation and does not yet reach the agentic case. Audience data from the underlying conference session show attendees already recognizing both the vulnerability and the gap in institutional guidance. Polled attendees most often named online quizzes as agent-completable, with discussion-based work close behind. Majorities in both listings were working without settled written guidance. The response defended here is design rather than detection: four principles for verified human presence, low-effort changes faculty can adopt now, and the assurance levers assessment professionals already operate.
Stavros P. Hadjisolomou, R. El-Haddad· Intersection: A Journal at t...· 0 citations
EduMind is introduced, a unified tutoring and assessment platform designed around a dual-track evaluation model that demonstrates how assessment and tutoring can be unified into a seamless workflow, and remained operationally stable throughout all testing phases.
Dhyan Gowda, M. Aruna, P. Prasad et al.· International Journal of Sci...· 0 citations
Large language models can produce fluent tutoring dialogue, but their educational use remains limited by weak grounding, uneven pedagogical control, and the risk of unsupported content. This study compared two controlled algebra tutoring workflows built on the same foundation model, Gemini 2.5 Flash.
The first condition was a prompt-only tutor that used the immediate task and scripted follow-up only. The second was an integrated pedagogical RAG tutor that combined a curated algebra corpus, learner-state variables, and a pedagogical response policy. In total, 24 two-turn algebra episodes were evaluated. No students were recruited, no classroom intervention was conducted, and no learning outcomes were measured. Eight expert reviewers scored anonymized response pairs for solution accuracy, conceptual support, instructional quality, instructional helpfulness, and unsupported content.
The pedagogical RAG condition received higher task-level expert ratings for conceptual support (M = 4.69 vs. 4.31), instructional quality (M = 4.61 vs. 4.29), and instructional helpfulness (M = 4.65 vs. 4.39). It also achieved a higher pedagogical-quality composite score (M = 4.65 vs. 4.33; Wilcoxon
p
= .00024; rank-biserial r = .86). Solution accuracy was high in both conditions and showed only a small descriptive difference (M = 4.84 vs. 4.74;
p
= .0769). However, unsupported content flags were more frequent in the RAG condition, occurring in 13 of 192 rating cells compared with one of 192 rating cells in the baseline condition. Fisher's exact test yielded
p
= .0015.
These findings do not establish learning gains, classroom efficacy, or the separate causal value of retrieval alone. They indicate that an integrated learner-state-aware RAG architecture can receive stronger expert ratings for pedagogical text quality on a fixed task set while also introducing trace leakage and evidence-boundary risks that must be controlled before deployment
Jude A. Adenuga· Frontiers in Education· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.