Human-centered GenAI feedback design in higher education: a multisite experiment on direct, reflective, and hybrid approaches to scientific argumentation
Jul 2026· International Journal of Educational Technology in Higher Education· Vol 23· 2 citations· 22 references
TL;DR
The findings suggest that the educational value of GenAI in higher education may depend less on AI access per se than on whether feedback environments preserve student agency, evaluative judgment, and ownership during revision.
Abstract
Generative artificial intelligence (GenAI) is increasingly used for feedback in higher education, yet evidence remains limited on how alternative human–AI feedback designs shape learning processes and durable outcomes. This study addresses that gap through a multisite, cluster-randomized, longitudinal field experiment comparing four feedback designs in introductory university science courses: peer feedback only, direct GenAI-supported feedback, reflective GenAI-supported feedback, and a hybrid design combining self-evaluation, peer feedback, and GenAI critique. The analytic sample comprised 1,176 first-year undergraduate students from 48 course sections across four universities and three science domains. Primary and secondary outcomes were argument-quality gain on four shared rubric dimensions—claim quality, evidence relevance and sufficiency, coherence of reasoning, and treatment of limitations or alternative explanations—conceptual learning, and delayed AI-free transfer; feedback uptake and self-regulated learning during revision were modeled as process mediators. Direct GenAI-supported feedback improved immediate argument-quality gain relative to peer feedback, whereas reflective and hybrid designs produced stronger feedback uptake and self-regulated learning. The hybrid condition yielded the highest adjusted mean for immediate argument-quality gain and showed the clearest advantage on conceptual learning; the reflective condition showed a positive but non-significant adjusted contrast on conceptual learning relative to direct GenAI-supported feedback. Both reflective and hybrid conditions outperformed direct GenAI-supported feedback on delayed AI-free transfer. Multilevel mediation analyses indicated that feedback uptake and self-regulated learning partially explained these advantages. By comparing four feedback designs, modeling revision processes, and assessing delayed AI-free transfer in a multisite field experiment, the findings suggest that the educational value of GenAI in higher education may depend less on AI access per se than on whether feedback environments preserve student agency, evaluative judgment, and ownership during revision.
This study aims to systematically synthesize empirical evidence on the effects of GenAI-mediated feedback on EFL writing, students’ uptake of such feedback, and the ethical issues associated with its use.
Following PRISMA, this review synthesized English-language, peer-reviewed studies published between 2023 and 2025. Searches were conducted in Scopus, ERIC and selected publisher platforms. Two reviewers independently screened records and extracted data using a piloted form, while risk of bias was assessed using RoB 2, ROBINS-I and JBI tools. Random-effects meta-analyses were conducted when at least three comparable effect sizes were available; otherwise, the evidence was synthesized narratively with harvest plots.
The review shows that GenAI-mediated feedback is most effective when implemented as a dialogic, multi-draft and criteria-aligned practice moderated by teachers. Students tend to adopt local corrections more readily than global revisions unless tasks explicitly require justification and verification. Human–AI agreement appears sufficient for low-stakes screening and formative diagnosis, but construct coverage remains uneven and subgroup differences emerge in some contexts, indicating the need for human oversight, transparent criteria and routine bias audits.
The evidence base remains fragmented across contexts, designs and reporting practices, which limits comparability and cumulative interpretation. Future research should standardize reporting, examine fairness across learner groups and evaluate the long-term pedagogical effects of GenAI-mediated feedback at scale.
EFL programs should develop prompt templates that elicit explanations and examples, incorporate verification routines, use public rubrics and exemplars and position teacher-curated co-feedback as a central component of implementation. Clear boundaries are also needed between formative and summative uses, together with disclosure requirements for AI-supported feedback.
This review provides a timely synthesis of recent evidence on GenAI-mediated feedback in EFL writing by integrating learning outcomes, uptake patterns and assessment fairness considerations within a single analytical framework.
U. Sulistiyo, Desti Angraini, Habib Maulana Fikri· Learning Futures and Emergin...· 0 citations
Generative artificial intelligence (GenAI) is increasingly used to explain concepts, summarize literature, suggest methods, generate arguments, and revise academic writing. These functions may reduce entry barriers to complex learning tasks while reshaping the cognitive and metacognitive work through which learners develop understanding. Doctoral research learning provides a high-complexity context for examining this issue because students must move from understanding existing knowledge to producing, justifying, and defending original claims.
This study examines how GenAI-mediated cognitive offloading redistributes cognitive work in doctoral research learning, how it may support or weaken learner-controlled engagement, and how learners recalibrate and reconstruct AI-supported outputs into defensible understanding and research judgment.
A grounded theory design was adopted. After four pilot interviews were used to refine the interview protocol, formal data were collected from 28 semi-structured interviews, 42 voluntarily provided AI-use artifacts, and 12 stimulated-recall discussions. Open coding, axial coding, selective coding, constant comparison, theoretical sampling, and theoretical saturation testing were used to construct a mechanism model. Interview transcripts were coded in Chinese to preserve participants' original meanings, and representative quotations were translated into English after analysis.
The analysis generated a five-part mechanism consisting of task-entry scaffolding, modular generation, metacognitive calibration crisis, critical reconstruction, and research accountability. Episodes reflected more learner-controlled engagement when GenAI reduced entry difficulty while preserving students' reading, verification, and reconstruction. Substitution risk appeared when AI-generated frameworks, method suggestions, thesis structures, or texts appeared coherent before learners had developed corresponding understanding. Calibration crisis appeared when recognition was mistaken for understanding, fluency for mastery, and coherence for justification. Critical reconstruction addressed this risk through source verification, self-explanation, contextual adaptation, and counterargument generation.
Learner-controlled engagement was most evident in episodes where doctoral students transformed AI-generated outputs into owned understanding, calibrated judgment, and defensible research decisions. These episodes suggest that accountable reconstruction helps preserve learners' responsibility for verification, reasoning, and final judgment in GenAI-supported doctoral learning.
Ling-Yun Gao, Feng Zhang· Frontiers in Psychology· 0 citations
Feedback is increasingly conceptualised as a dialogic and socio-culturally embedded process; however, little is known about how feedback practices operate across structural levels in laboratory-based higher education contexts. This study investigates how feedback is enacted across micro-, meso-, and macro-level dimensions of feedback culture in undergraduate chemistry laboratory courses within a Kazakhstani university. A qualitative case study design was employed, combining naturalistic classroom observations and semi-structured interviews. Data were collected from three laboratory instructors and 40 undergraduate students across twelve observed laboratory sessions (approximately 24 contact hours), supplemented by interviews with 3 instructors and 3 students. Observation and interview data were analysed using a hybrid deductive–inductive thematic approach guided by feedback literacy frameworks. Findings indicate that feedback was frequent but predominantly enacted at the micro level as immediate procedural correction supporting task completion and laboratory safety. Meso-level feedback processes such as reflection, feedforward, and continuity across sessions were weakly structured, while macro-level structural support for feedback was limited. These cross-level conditions produced a stable configuration conceptualised as procedural feedback lock-in, whereby corrective feedback becomes normalised despite pedagogical intentions toward developmental learning. The study contributes to feedback culture scholarship by extending feedback literacy research into laboratory-based STEM education and offering analytically transferable insights for strengthening feedback design in practice-intensive higher education environments.
M. Mataev, A. Çelik, Moldir Abdraimova et al.· Frontiers in Education· 0 citations
Examining a four-process GenAI cycle reveals two distinct metacognitive regulation styles: Exploratory-Simplification and Systematic-Methodical, which show that students use combinations of strategies across the AI-SRL cycle.
Maria A. Perifanou, Anastasios A. Economides· Journal of educational compu...· 0 citations
This research compares AI-generated and teacher feedback on student reflections in Japanese K-12 classrooms that adopt self-paced learning, focusing on feedback characteristics and learner receptivity. Two AI conditions were designed based on Hattie and Timperley’s four-level feedback model: Theory-AI, employing GPT-4.1-mini with this framework alone, and Context-AI, employing GPT-5.1 with an Intelligent Tutoring System (ITS) architecture that integrates learner and domain models. These were compared with teacher feedback (Teacher) in a within-subject design involving 110 students across four schools and five grade levels. Feedback characteristics (RQ1) were analyzed along three dimensions—volume, generation time, and content composition across the four levels (task, process, self-regulation, and self)—and their uniformity both between and within class groups. Learner receptivity (RQ2) was assessed via clarity, specificity, and empathy on a 5-point scale. For RQ1, both AI conditions exhibited higher proportions of process- and self-regulation-level feedback than teachers, whereas Theory-AI produced excessive self-level feedback. AI feedback was generated approximately ten times faster and exhibited substantially greater uniformity in both volume and content composition, whereas teacher feedback varied considerably between class groups (e.g., self-level inclusion: Cramér’s $V = .468$ ) and within the same class group (character-count SD: 9.4–37.1). For RQ2, no differences were observed in clarity; specificity ratings followed Context-AI > Theory-AI > Teacher, and empathy ratings followed Theory-AI > Context-AI > Teacher, with all pairwise differences significant. These findings indicate that ITS-integrated AI feedback offers an efficient, theoretically grounded complement to teacher feedback in K-12 education.
Fostering design thinking relies on instructional feedback, yet traditional instructor‐led methods face practical limitations. Although generative AI (GenAI) may offer promising solutions based on its advantages, existing studies predominantly characterise it as a performance‐enhancing task assistant rather than as a feedback‐providing pedagogical agent for the intrinsic growth of design thinking. The study compares the pedagogical effects on students' design thinking and students' perceptions of feedback systems by GenAI and human instructors. A within‐class randomised experimental design was conducted with 80 undergraduates. Results indicate no significant overall difference in learning gains, but reveal respective stage‐specific strengths: GenAI proved more effective during the ‘empathise’ stage, associated with superior timeliness and privacy, while human instructors excelled in the ‘prototype’ stage, valued for their contextual anchoring. Comparable effects yet complementary functions were observed in other stages, reinforcing GenAI's augmentative role instead of substitution. Findings collectively demonstrate that the effectiveness and strengths of feedback systems are not static and uniform but dynamically aligned with the learning demands across educational settings. Practical insights for designing hybrid intelligence feedback systems and theoretical implications were finally discussed.
Instructional feedback is crucial for fostering design thinking; however, traditional instructor‐led feedback is often constrained by issues of scalability, timeliness and consistency.
Generative AI (GenAI) shows promise to serve as an additional feedback agent, but its role has been explored more as a performance‐enhancing task assistant rather than a feedback‐providing pedagogical agent for the intrinsic growth of design thinking.
The study demonstrates that GenAI leads to greater learning gains in the ‘empathise’ stage, while human instructors are more effective in the ‘prototype’ stage.
The study underscores their comparable effect while serving distinct and complementary roles at ‘define’, ‘ideate’ and ‘reflect’ stages.
The study highlights their respective stage‐specific strengths and associates them with learning gains across the design thinking process.
The study provides practical insights for designing hybrid feedback systems that strategically orchestrate GenAI and human instructors to maximise pedagogical impact.
Educational policy should promote the design of hybrid feedback systems that strategically align the inherent characteristics of different feedback agents with the stage‐specific learning demands across various educational contexts.
Curriculum development should emphasise building students' relevant competencies (e.g., feedback literacy, AI competency, etc.), equipping them to seek, evaluate and synthesise input from multiple agents effectively.
Xiao Fei· British Journal of Education...· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.