Automated Writing Evaluation: Assessing Detection Performance, Feedback Accuracy, and Pedagogical Value
This thesis examines how automated writing evaluation (AWE) systems function as feedback providers in English-as-a-foreign-language (EFL) writing. Despite their widespread use in writing classrooms, AWE systems are often evaluated primarily in terms of overall accuracy. This thesis adopts a multi-dimensional evaluation framework that integrates error detection, feedback accuracy, and the pedagogical value of feedback to examine the per-formance of AWE systems. To operationalise this framework, fourteen AWE systems were analysed using a text. The text was written by the researcher, who has an EFL background, in his first year of undergraduate study. The researcher and his supervisors (with L1 English backgrounds) manually evaluated the text. System outputs were compared with the human evaluation across error types: spelling, grammar, diction, tone and style, and coherence and cohesion. The analysis examined overall error detection performance and feedback accura-cy, as well as the pedagogical value of feedback in terms of feedback types: direct, semi-direct, indirect, and metalinguistic. The findings reveal substantial variation in system performance within and between systems, with a consistent skew towards prioritising lower-level error detection. Most sys-tems demonstrated limited effectiveness in addressing higher-level writing errors. Skewedness can not only be found in error detection performance but also in the types of feedback generated by systems. Most systems relied heavily on direct corrections and ge-neric explanations, whereas semi-direct/indirect and metalinguistic feedback were less fre-quently generated and varied in quality. Although no system performed consistently well across all error types, premium Grammarly (2025) demonstrated the best overall error de-tection performance across error types and feedback quality among the fourteen AWE sys-tems evaluated in this thesis. Consequently, premium Grammarly (2025) was identified in this thesis as the most effective system based on the integrated evaluation of error detec-tion performance, feedback accuracy, and pedagogical affordances of feedback. These findings demonstrate that the effectiveness of AWE systems cannot be ade-quately understood through overall accuracy alone, thereby challenging accuracy-driven or score-driven approaches to automated writing evaluation. Instead, system performance must be interpreted in terms of error detection across error types and how feedback sup-ports learning. While AWE systems are effective in correcting lower-level errors, their lim-ited capacity to address higher-level writing problems and to generate pedagogically useful feedback constrain their role as learning resources. Moreover, these findings highlight the need for a broader evaluative perspective. Such an evaluative perspective could consider whether automated feedback should be embedded in a learning process and the extent to which automated feedback promotes learner engagement.