Skip to content
Review Open access

A Mobile System for Formative Assessment and Teacher-Reviewed Feedback

Aug 2026 · International Journal of Interactive Mobile Technologies (ijim) · 1 citation

TL;DR

Within the limits of this small-scale classroom evaluation, the findings suggest that rubric-guided LLM scoring may support formative assessment in mobile classrooms, provided teacher oversight is maintained for borderline cases.

Abstract

Open-ended short responses can reveal more about student understanding than selectedresponse items yet scoring them during live instruction is rarely feasible. This paper presents a bring-your-own-device (BYOD) mobile classroom system that combines rubric-guided large language model (LLM) scoring of short open-ended answers, real-time teacher analytics through a classroom dashboard, and evidence-based post-quiz feedback generation. The system includes a teacher-reviewed workflow in which instructors can inspect model outputs and, when needed, adjust scores or feedback during formative classroom use. In the reported evaluation, however, teacher override was not applied or analyzed; the agreement metrics compare the system’s initial rubric-guided outputs with expert reference scores. The system was evaluated in three classroom groups (N = 60 students; 480 question-level scoring cases). Three expert raters participated, each assigned to one classroom group, so each response received one expert reference score. In this test, system scores showed strong but not complete agreement with expert reference scores (QWK = 0.887; MAE = 0.089; RMSE = 0.121). Mastery-label accuracy reached 78.3%. Expert raters gave positive scores to the generated feedback reports for groundedness (M = 4.78), specificity (M = 4.37), and actionability (M = 4.57) on a 1–5 scale. Within the limits of this small-scale classroom evaluation, the findings suggest that rubric-guided LLM scoring may support formative assessment in mobile classrooms, provided teacher oversight is maintained for borderline cases.

Read PDF

Similar papers

Review Open access Aug 2026

The COM Essay Assessor

The development and calibration of the COM Essay Assessor is presented, a rubric-based generative artificial intelligence (GenAI) tool designed to support formative feedback while retaining instructor oversight and reflects on the opportunities and challenges of integrating GenAI into large writing programs.

Juhi Bansal · 0 citations
Review Open access 2026

From Whole-Class Replay to Student-Controlled Listening Support: An Action Research Study of Interactive Listening Webpages in EAP

This action research study was conducted in a Year 1 Foundation EAP listening classroom at Xi’an Jiaotong-Liverpool University. It addressed the problem that teacher-controlled whole-class replay cannot fully meet individual needs in a mixed-ability class. An interactive HTML listening webpage was introduced so that students could submit answers on their own devices, receive immediate feedback, replay question-specific audio clips, and consult segmented transcripts as delayed support. Data were collected from 11 survey responses and three semi-structured interviews, and were analysed through descriptive statistics and thematic analysis. The findings suggest that students valued the webpage because it enabled self-paced review, targeted replay, and more private checking of mistakes. However, the tool should be accompanied by teacher guidance on listening strategies, transcript use, and device management. The study offers practical implications for differentiated and reflective listening support in EAP classrooms.

仕峥 刘 · 0 citations
Open access Aug 2026

AI-Assisted Web-Based Corrective Feedback To Improve Senior High School Students’ Listening Comprehension

Listening receives limited class time at the senior high school level, and large classes rarely allow teachers to give personal written feedback on students’ open-ended answers. Web-based listening tools and AI scoring systems both exist, but the two are seldom integrated so that learners receive immediate corrective feedback on free-text listening responses. This study addressed that gap by developing and testing a web-based learning system assisted by artificial intelligence (AI), and it examined both the feasibility and the effectiveness of the product. The Research and Development method was applied through the Integrative Learning Design Framework across three phases, namely exploration, enactment, and evaluation. The product is a website named Ear Up! that delivers CEFR-graded audio lessons, collects essay answers, and uses an OpenAI GPT-4o model to return automated corrective feedback and next-lesson recommendations. The system was trialed with 28 tenth-grade students. Expert and user evaluation produced a mean feasibility score of 4.67 out of 5, in the very feasible band. Under a one-group pretest-posttest design, listening for details improved at a medium level (N-gain = 0.67) and listening for inference at a low level (N-gain = 0.25), while classical mastery rose from 29.41% to 82.35% and from 32.35% to 67.65%. The findings indicate that automated corrective feedback can extend listening practice beyond limited class time, yet higher-order inference still depends on teacher mediation. AI-assisted corrective feedback is therefore best positioned as a teacher-supervised supplement in future language learning.

Marie Louise Catherine Widyananda, M. K. Wirasti, Dwi Kusumawardani · 0 citations
Open access Jul 2026

Screencast versus Text-Based Feedback in L2 Academic Writing: Revision Uptake and Student Perceptions

The rapid expansion of technology-mediated feedback modes has prompted growing interest in screencast feedback as an alternative to text-based electronic feedback in L2 writing instruction. This study compares the effects of screencast feedback and text-based electronic feedback on revision uptake and explores students’ perceptions of screencast feedback in a content-based academic writing course. The participants were 41 pre-service English language teachers enrolled at a state university in Türkiye. Over one semester, the experimental group received indirect instructor feedback on their second drafts through screencast video recordings, while the control group received equivalent feedback as marginal comments in Microsoft Word. Quantitative data were drawn from a randomly selected subsample of 20 students (10 per group), resulting in 40 final drafts per group for analysis. Revision uptake was measured by calculating correction scores from final drafts, and group differences were analysed using Mann-Whitney U tests. In addition, six randomly selected participants from the experimental group took part in a semi-structured focus group interview to explore their perceptions of the screencast feedback process. The results showed that students in the screencast feedback group achieved significantly higher correction scores at the micro, macro, and overall levels. Qualitatively, students perceived screencast feedback as improving their writing, enhancing their engagement with the feedback process, and facilitating their revision strategies, while also demonstrating critical awareness of the constraints imposed by time limitations on feedback delivery. These findings contribute to the growing evidence base for screencast feedback in L2 writing and highlight its potential as an effective and engaging feedback mode in EFL pre-service teacher education contexts.

G. Kurt, Aslıhan Karabulut · 0 citations
Open access Sep 2026

Formative Feedback Beyond Predefined Options: Eliciting Student-Generated Concepts in Large Classes by Adding Very Short Answers to an Audience Response System

Within appropriate instructional frameworks, Audience Response Systems (ARS) can support active learning, retrieval practice, and feedback. However, they typically rely on selecting from predefined options, which limits active answer generation and insight into students’ reasoning. In this feasibility study, we integrated very short answers (VSA) into a hybrid ARS and examined whether free-text responses can be reliably aggregated in large classes, whether they extend student concepts beyond predefined options, and how students perceive the approach. First-semester medical students (n = 326) answered 26 questions in alternating multiple-choice (MC) and VSA formats, yielding 11,561 analyzed responses. Responses were automatically aggregated, and the ten largest aggregates per question were displayed in real time to guide discussion and feedback, capturing 74.6% of all responses and 40.3% of VSA responses. Students submitted more responses per answered question in MC than in VSA (mean difference = 0.35, 95% CI [0.31, 0.38], p < 0.001). Post hoc consolidation of conceptually identical expressions yielded 7.6 distinct concepts per question within the Top10 Aggregates, with VSA responses contributing concepts beyond the predefined MC options. Respondents perceived the approach as enhancing understanding and reducing anxiety. By eliciting student-generated concepts and reasoning, including unanticipated ones, VSA integration broadens formative feedback in ARS.

Aliya Tuktarova, M. Grasl, I. Volf · 0 citations
Open access Aug 2026

Student’s Perception on The Implementation of Project-Based Learning

This study investigated the implementation of the Project-Based Learning (PjBL) model integrated with technology and speaking skills, as well as students’ perceptions of its implementation, among eleventh-grade students in the Hospitality Accommodation program at SMK Negeri 1 Seririt. Using a sequential exploratory mixed-methods design, qualitative data were collected through structured classroom observations based on the Gold Standard PjBL framework, while quantitative data were obtained using a closed-ended Likert-scale questionnaire based on a perception framework and administered to 135 students. The qualitative findings revealed that the English teacher systematically implemented PjBL by anchoring projects in authentic hospitality-related challenges and differentiating activities according to students’ English proficiency levels (low, moderate, and high). Students engaged in sustained inquiry using digital resources, managed their learning within a structured four-meeting framework, participated in peer critique and revision, and produced multimedia outputs such as PowerPoint presentations, digital infographics, and practical simulation videos presented in the classroom. Furthermore, the quantitative analysis showed that students had a “Very Positive” perception of PjBL implementation, with an overall mean score of 131.29. All core perceptual dimensions—Interpretation (4.40), Organization (4.39), and Selective Attention (4.33)—consistently fell within the very positive category. These findings indicate that PjBL effectively bridges theoretical instruction and practical professional communication skills while fostering cognitive engagement, metacognitive independence, and oral proficiency relevant to the hospitality industry.

Made Dina Martini Putri, Putu Adi Krisna Juniarta, Dewa Ayu Eka Agustini · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.