It is suggested that current efforts to mitigate mode collapse are insufficient for open-ended educational generation, and new training and data collection strategies to support pedagogical diversity are needed.
While various organizations now actively encourage LLM use in classrooms, we still lack rigorous, systematic evaluations of how well these models actually perform the fundamental tasks of language pedagogy. This paper examines whether state-of-the-art LLMs can deliver the kind of corrective feedback and methodological explanations that language learners need. The study tests multiple large language models on their ability to identify, correct, and explain common learner mistakes in English, by systematically varying model parameters to investigate how these technical adjustments affect output quality, pedagogical clarity, and consistency, along with using retrieval-augmented generation to query methodological data. The evaluation employs automated metrics (GLEU, BERTScore) but also human expert judgments to capture dimensions that purely computational measures miss: linguistic nuance, cultural sensitivity, and instructional appropriateness. While models demonstrate impressive surface-level correction abilities, their explanations often lack the terminological and domain knowledge that effective language teaching requires, suggesting that current enthusiasm for AI-assisted language learning may be outpacing our understanding of these systems'actual pedagogical competence.
Kristina Šekrst, A. Kovačič· Oxford Intersections: AI in...· 0 citations
For teachers to effectively use large-language-model(LLM)-based ratings in the formative or summative assessment of texts, it is essential to ensure that such ratings can assess student writing in a valid and reliable manner. This study investigates whether a validated human text-rating procedure (benchmark rating) can be replicated by an LLM-based rating procedure. We tested the replication with two genres of elementary school students’ text—narrative and instructive—using nine LLMs from three providers (OpenAI, Anthropic, Mistral). Each LLM generated three independent scores per text via structured, benchmark-aligned prompts that were then aggregated into a consensus score. Results showed that intrarater reliability was high to excellent, ICC(3, k) ≈ .68–.97, and alignment with human ratings ranged from moderate to strong, ICC(3, 1) ≈ .47–.85, with larger models consistently outperforming smaller ones. Systematic bias patterns emerged, varying by model and genre, indicating a need for calibration. Increasing output token windows and reducing temperature parameters mitigated truncation and schema-related failures. Although LLM-based benchmark ratings can approximate expert judgments and reduce the need for labor-intensive human triple coding, limitations remain regarding cost (for larger models), genre- and task specificity, and sensitivity to text presentation and student grade level—factors that constrain immediate classroom use, particularly for formative feedback.
Generative AI can produce fluent prose before students have worked through the linguistic choices that writing courses are intended to teach. This study examined how one university instructor responded to that problem in a second-year English-major writing course. Forty-two students (18 men and 24 women; mean age = 20.3 years) participated in a 16-week concurrent embedded mixed-methods study. The instructor compared AI-probability estimates from GPTZero, Originality.AI, and Turnitin, returned passage-level feedback, and taught four sessions on making context-sensitive lexical choices without asking AI to rewrite students' work. Questionnaires, interviews with six students, three sets of writing samples, and course scores supplied the data. Across Weeks 5, 10, and 15, the mean estimated share of AI-generated text fell from 51.2% to 35.8% and then to 23.6%. Overall AI-tool use also declined, while mean writing scores rose from 72.4 to 85.9. Students most often associated AI use with difficulty expressing ideas accurately, limited time, and pressure to obtain higher grades. By the end of the semester, more students distinguished grammar assistance from having AI compose a paper. Because the study involved one class and no control group, these changes should be read as associations rather than proof of a causal effect. Even so, the findings illustrate how transparent checking, individualized revision, and lexical instruction can be combined in a classroom response to AI-dependent writing.
Yue-Han Hou· Communications in Humanities...· 0 citations
Large language models are increasingly deployed in education as tutors, teaching assistants, and content generators. These roles place demands that ordinary question answering does not: a usable education-facing model is supposed to be accurate, safe under sensitive prompts, instructionally useful, and aligned with pedagogical goals at the same time. Existing benchmarks evaluate these requirements largely in isolation, so none assesses education-facing suitability as an integrated profile. We introduce ELBench, the first benchmark to evaluate all four requirements (General Capability, Safety and Trustworthiness, Basic Education, and High-Level Cultivation) on the same models under a common protocol, combining curated public sources with newly synthesized safety and cultivation data. We evaluate nine models, seven frontier general-purpose systems and two education-specialized variants, and report three findings. First, module-level profiles are more informative than a single aggregate: the top six models are statistically indistinguishable on overall score, yet their module leaders differ substantially, and safety is anti-correlated with practical teaching (r = -0.83). Second, the Chinese-developed models lead the safety module, the most discriminative in the suite; this advantage is largest on region-specific normative content and narrows, but does not vanish, on universal-harm content. Third, the two education-specialized models lead neither education module, and on High-Level Cultivation all models share a systematic blind spot: on the structured judgment task they converge on the same non-reference option, favoring pedagogical style over fit to the stated goal, so the module scores uniformly low and does not separate models. This raises, but does not resolve, whether domain post-training keeps pace with frontier systems on education tasks.
Yi-Lin Jiang, Xiao-Rong Zhu, Fei Tan et al.· 0 citations
The increasing prevalence of artificial intelligence (AI), in particular large language models (LLMs), is transforming how scientific texts are produced and consumed. A growing share of online content is at least partially generated by AI, often without readers’ awareness, raising the question of whether students can recognise such texts and under what conditions. In this study, 372 high school students evaluated short physics texts and indicated whether each had been written by a human author or by ChatGPT. The texts were organised into tasks that varied in topic and other task characteristics, including cognitive demand. Overall, students struggled to identify AI-generated texts, performing only slightly above chance. Focusing on task-level features, we analyse how recognition accuracy varies across texts and discuss which characteristics make AI-generated physics texts more or less distinguishable from human-written ones, with implications for fostering AI literacy in physics education.
Péter Kosztyó, Márton Burkovics, Péter Jenei· Journal of Physics, Conferen...· 0 citations
Writing proficiency manifests in how students develop content, organize ideas, choose words, and use language. Despite growing interest in LLM-based student simulation, whether LLMs can reproduce such multidimensional variation in extended writing remains largely unexplored. In this work, we explore if language models can realistically simulate student writing, and introduce SWIM, a task that formulates Student Writing sIMulation as proficiency-conditioned essay generation. We evaluate prompting, supervised fine-tuning (SFT), and reinforcement learning (RL) methods for writing simulation using automated essay scoring as a measure of profile alignment. Extensive experiments reveal that prompting provides limited proficiency control, even for strong proprietary LLMs with rubric-grounded strategies. In particular, while models can adjust content-oriented traits, they struggle to reproduce the lexical, grammatical, and organizational variation in different proficiency levels. SFT substantially improves alignment, while RL with the proposed proficiency-alignment reward yields further gains across all writing traits and essay prompts. Our findings suggest that explicit supervision enables substantially stronger profile alignment than prompting alone, while authentic low-proficiency writing remains challenging to reproduce.
Heejin Do, Jakub Kontak, Mrinmaya Sachan· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.