Aug 2026· The International Journal of Technologies in Learning· 0 citations
TL;DR
The development and calibration of the COM Essay Assessor is presented, a rubric-based generative artificial intelligence (GenAI) tool designed to support formative feedback while retaining instructor oversight and reflects on the opportunities and challenges of integrating GenAI into large writing programs.
Abstract
Providing timely and actionable feedback on student writing is a known challenge in large English as a Second Language (ESL) classrooms, where instructor workload often limits the depth and consistency of feedback. This article presents the development and calibration of the COM Essay Assessor, a rubric-based generative artificial intelligence (GenAI) tool designed to support formative feedback while retaining instructor oversight. Grounded in social constructivist theory and formative assessment research, the tool was calibrated using archival student essays and faculty feedback to reflect course-specific evaluation practices. Rather than functioning as an autonomous evaluator, the tool operates within a human-in-the-loop workflow in which instructors review and authorize AI-generated feedback before it reaches students. The article situates this approach within broader work on automated writing evaluation and AI-supported learning and reflects on the opportunities and challenges of integrating GenAI into large writing programs. It argues that rubric-aligned, human-mediated systems can help sustain feedback processes in resource-constrained contexts while preserving pedagogical intent and instructor judgment.
Analysis of reflective essay-feedback-appraisal instances from 283 Estonian bachelor students across one semester contributes descriptive classroom evidence on integration of AI feedback - a fast and scalable way to provide immediate writing advice, but not a self-contained route to better reflection.
Andres Karjus, Janika Leoste, Tiia Õun· arXiv.org· 0 citations
This study examines feedback in English as a Foreign Language (EFL) writing contexts, focusing on written corrective feedback (WCF). Large language models (LLMs) can provide WCF at scale, but aligning them with pedagogical best practices remains an ongoing challenge. WCF meeting criteria like factuality or relevance may still be unsuitable for learning contexts, highlighting the need for extrinsic evaluation based on the learner's perspective. We deployed WCF systems in a university-level EFL class with nearly 2,000 students, collecting over 20,000 drafts. We evaluated the generated WCF from two perspectives: intrinsic evaluation by experienced English teachers using a rubric, and extrinsic evaluation via student feedback and engagement metrics. Results revealed low alignment between teacher expert ratings and student feedback. These findings suggest that traditional expert evaluation alone may not fully capture WCF's usability or helpfulness from the learner's perspective, highlighting the importance of learner-centered evaluation frameworks for AI-based applications in language education.
Steven Coyne, Diana Galván-Sosa, Ryan Spring et al.· arXiv.org· 0 citations
The findings suggest that AI-generated feedback supported targeted revision when it is accessible, interpretable, and aligned with classroom assessment criteria.
A. Tzirides, Michele Galla, B. Cope et al.· Ubiquitous Learning An Inter...· 0 citations
ABSTRACT Generative artificial intelligence (GenAI) tools are an increasingly common resource used in the classroom and writing process. The landscape of available GenAI tools is rapidly evolving, so having a systematic and straightforward way to evaluate new tools for incorporation into the classroom is key. For example, in science writing education, GenAI tools can be used as a supplement to instructor feedback on student writing, allowing an additional opportunity for critique and revision by the student. Here, we describe a rubric we developed to enable instructors to assess differences between feedback provided by GenAI models based on five key areas: (i) accuracy, (ii) constructiveness, (iii) clarity and readability, (iv) recognition of strengths, and (v) original text. We show data comparing the feedback provided by multiple GenAI models on student work based on course guidelines. This rubric can serve as a resource for other instructors interested in evaluating various GenAI models for use as scientific writing feedback supplements in their own classrooms.
Daniel R. Rankins, Erick N. Tran, Valerie T. La et al.· Journal of Microbiology & Bi...· 0 citations
Within the limits of this small-scale classroom evaluation, the findings suggest that rubric-guided LLM scoring may support formative assessment in mobile classrooms, provided teacher oversight is maintained for borderline cases.
Abylay Yerniyazov, Bakyt Bakayeva, Zhalgasbek Iztayev et al.· International Journal of Int...· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.