Aug 2026· Oxford Intersections: AI in Society· 0 citations· 44 references
Computer Science
Abstract
While various organizations now actively encourage LLM use in classrooms, we still lack rigorous, systematic evaluations of how well these models actually perform the fundamental tasks of language pedagogy. This paper examines whether state-of-the-art LLMs can deliver the kind of corrective feedback and methodological explanations that language learners need. The study tests multiple large language models on their ability to identify, correct, and explain common learner mistakes in English, by systematically varying model parameters to investigate how these technical adjustments affect output quality, pedagogical clarity, and consistency, along with using retrieval-augmented generation to query methodological data. The evaluation employs automated metrics (GLEU, BERTScore) but also human expert judgments to capture dimensions that purely computational measures miss: linguistic nuance, cultural sensitivity, and instructional appropriateness. While models demonstrate impressive surface-level correction abilities, their explanations often lack the terminological and domain knowledge that effective language teaching requires, suggesting that current enthusiasm for AI-assisted language learning may be outpacing our understanding of these systems'actual pedagogical competence.
While AI is often viewed as a threat to academic integrity, this paper proposes tools like Grammarly as sophisticated pedagogical allies bridging Natural Language Processing (NLP) theory and classroom practice. Based on a systematic review of 15 studies (2018–2025), the research introduces two implementable frameworks: the AI-Assisted Diagnostic Assessment model for efficient feedback and the 5-Step Structured Writing Protocol for student scaffolding. The analysis centers on the "5 Cs"—Coherence, Correctness, Clarity, Conciseness, and Consistency. The study argues that delegating mechanical corrections to automated systems enables educators to focus on higher-order skills, including critical thinking and authorial voice. This shifts the instructional dynamic toward a "trialogue" (Teacher, Student, AI), fostering a scaffolded partnership crucial for developing linguistic autonomy and fluency in both L1 and L2 contexts.
Sofia Tsakalidou· International Journal of Edu...· 0 citations
It is suggested that current efforts to mitigate mode collapse are insufficient for open-ended educational generation, and new training and data collection strategies to support pedagogical diversity are needed.
Mei Tan, Hyunji Nam, Dorottya Demszky· 0 citations
English writing proficiency is a fundamental academic competency and a key indicator of students' language mastery in higher education. Theoretically, this complexity is grounded in the Complexity, Accuracy, and Fluency (CAF) framework, with the accuracy aspect specifically emphasizing adherence to grammatical and spelling rules. However, the manual assessment processes typically employed by instructors are often hindered by heavy technical workloads, high operational costs, and potential inter-rater inconsistency, ultimately limiting the frequency and depth of formative feedback provided to students. While commercial Automated Writing Evaluation (AWE) tools offer a partial solution, many operate as opaque "black box" systems reliant on proprietary third-party APIs, raising concerns regarding data reliability, institutional privacy, and a lack of pedagogical transparency for learners. This research seeks to address these issues by developing a transparent, self-hosted writing assessment prototype using the Design Science Research (DSR) framework. The system is built on the Laravel 12 framework—chosen for its modularity and security—and integrates the open-source LanguageTool API to provide rule-based feedback on spelling and writing style. The prototype's performance was rigorously evaluated against a "gold standard" established by independent expert raters using a dataset of authentic student essays. Validation results demonstrate high reliability, with the system achieving significant accuracy in spelling and grammar detection. Furthermore, the system exhibits high technical efficiency with rapid response times. This research aims to produce a digital solution capable of further development that bridges the gap between traditional assessment methods and modern educational technology; the solution is also expected to effectively alleviate the technical workload of instructors while empowering students to become more independent through real-time instructional feedback based on language usage standards.
Anwar Hilman, Vivi Ayu Lestari, J. Sihombing et al.· bit-Tech· 0 citations
SocraticTrap-CS is introduced, a publicly available benchmark that probes the capacity of open-weight LLMs to generate strategic misconceptions on demand and reframes the evaluation of educational LLMs around pedagogical trustworthiness rather than factual correctness alone.
Marijela Miličević, Mia Rovis, Ratomir Karlović et al.· Information· 0 citations
Critical thinking is a fundamental skill that helps learners move beyond simple memorization. One way to develop this skill is through high-order questioning. However, crafting such questions remains a challenge for educators, and classroom practices tend to rely on low-order questions. Large Language Models have demonstrated strong capabilities in generating high-order questions, especially when guided by prompts based on Bloom's Taxonomy. Yet, existing research has largely centered on this framework and focused only on English. This study addresses these gaps by introducing prompts grounded in two alternative frameworks: Claim-Evidence-Reasoning and Divergent Questioning within a multilingual context using Basque, Spanish, and English. Results indicate that while both an open-source and a proprietary model rather effectively generate questions in all three languages, only about half of the answerable questions are recognized by teachers as high-order. A positive finding is that the alternative frameworks produce structurally and conceptually varied questions, suggesting they could complement each other and provide viable alternatives to Bloom's Taxonomy.
Suna-cSeyma Uccar, Itziar Aldabe, Nora Aranberri et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.