Skip to content
Preprint

Consistently Good vs. Occasionally Great: A Rubric for Open-Ended Feedback Quality from Humans and Machines

Aug 2026 · 0 citations · 44 references
Computer Science

TL;DR

This paper focuses on feedback for open-ended short answer questions in introductory programming, with the goal of nudging students toward success on reattempts without revealing the correct answer, and develops a five-criteria rubric grounded in educational literature for evaluating feedback quality.

Abstract

Providing high-quality feedback on student work is essential for learning, yet delivering such feedback at scale remains challenging. In this paper, we focus on feedback for open-ended short answer questions in introductory programming, with the goal of nudging students toward success on reattempts without revealing the correct answer. We develop a five-criteria rubric grounded in educational literature for evaluating feedback quality: (1) acknowledging correct portions of the student answer, (2) identifying at least one flaw (if present), (3) providing actionable guidance for improvement, (4) maintaining appropriate concealment of the answer, and (5) using an appropriate conversational tone. Using this rubric, we compare feedback generated by a frontier LLM (OpenAI o1) to feedback from nine teaching assistants across 90 student responses, with three researchers and an LLM independently scoring all feedback. Our results show that while one TA often produced the best feedback, the LLM demonstrated consistently higher average performance than TAs, as evaluated by humans. However, we also uncover significant self-preference bias when using LLMs to evaluate feedback quality: the LLM systematically rated its own outputs higher than human experts did. This bias, which research suggests persists even in cross-model evaluation, raises important methodological concerns for researchers employing LLM-based evaluation. We provide detailed characterization of both TA and LLM performance, analyze sources of variance in TA feedback quality, and discuss implications for deploying LLM-generated feedback in educational settings.

View source

Similar papers

Review Jul 2026

Computer-generated QIS tutorial feedback is valued by students, but does not replicate in-class collaboration

While the computer-generated feedback was broadly considered useful by students, student engagement patterns were markedly different in the solo setting, with students demonstrating reluctance to use the interface's built-in help features and tending to internalize failure in unproductive ways counter to the intention of a formative learning environment.

J. C. Meyer, S. Pollock, Bethany R. Wilcox et al. · 0 citations
Open access Jul 2026

How Well Does AI-Generated Feedback Work? Intrinsic and Extrinsic Evaluation across more than 20,000 EFL Essay Drafts

This study examines feedback in English as a Foreign Language (EFL) writing contexts, focusing on written corrective feedback (WCF). Large language models (LLMs) can provide WCF at scale, but aligning them with pedagogical best practices remains an ongoing challenge. WCF meeting criteria like factuality or relevance may still be unsuitable for learning contexts, highlighting the need for extrinsic evaluation based on the learner's perspective. We deployed WCF systems in a university-level EFL class with nearly 2,000 students, collecting over 20,000 drafts. We evaluated the generated WCF from two perspectives: intrinsic evaluation by experienced English teachers using a rubric, and extrinsic evaluation via student feedback and engagement metrics. Results revealed low alignment between teacher expert ratings and student feedback. These findings suggest that traditional expert evaluation alone may not fully capture WCF's usability or helpfulness from the learner's perspective, highlighting the importance of learner-centered evaluation frameworks for AI-based applications in language education.

Steven Coyne, Diana Galván-Sosa, Ryan Spring et al. · 0 citations
Review Open access Aug 2026

The COM Essay Assessor

The development and calibration of the COM Essay Assessor is presented, a rubric-based generative artificial intelligence (GenAI) tool designed to support formative feedback while retaining instructor oversight and reflects on the opportunities and challenges of integrating GenAI into large writing programs.

Juhi Bansal · 0 citations
Open access Aug 2026

A rubric to assess generative AI-based feedback on student writing assignments

ABSTRACT Generative artificial intelligence (GenAI) tools are an increasingly common resource used in the classroom and writing process. The landscape of available GenAI tools is rapidly evolving, so having a systematic and straightforward way to evaluate new tools for incorporation into the classroom is key. For example, in science writing education, GenAI tools can be used as a supplement to instructor feedback on student writing, allowing an additional opportunity for critique and revision by the student. Here, we describe a rubric we developed to enable instructors to assess differences between feedback provided by GenAI models based on five key areas: (i) accuracy, (ii) constructiveness, (iii) clarity and readability, (iv) recognition of strengths, and (v) original text. We show data comparing the feedback provided by multiple GenAI models on student work based on course guidelines. This rubric can serve as a resource for other instructors interested in evaluating various GenAI models for use as scientific writing feedback supplements in their own classrooms.

Daniel R. Rankins, Erick N. Tran, Valerie T. La et al. · 0 citations
Jul 2026

Student Evaluation of Repeated AI Feedback Across a Semester of Writing

Analysis of reflective essay-feedback-appraisal instances from 283 Estonian bachelor students across one semester contributes descriptive classroom evidence on integration of AI feedback - a fast and scalable way to provide immediate writing advice, but not a self-contained route to better reflection.

Andres Karjus, Janika Leoste, Tiia Õun · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.