Skip to content

Author

Xiaoliang Zhou

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

Evaluating AI-Based Automated Essay Scoring Through Signal Detection Theory: Beyond Aggregate Agreement Metrics.

The rapid adoption of Large Language Models (LLMs) in educational assessment has reshaped scoring practices, yet evaluation remains tethered to aggregate reliability metrics like Quadratic Weighted Kappa, which obscure discrimination and rater effects. This study applies Signal Detection Theory to evaluate eight state-of-the-art LLMs (including Claude 3.5 Haiku, DeepSeek-V3, Gemini 3 Flash, GPT-4o, and Grok 4.1) against expert human raters across 1,726 essays. By decoupling discrimination from response criteria, I provide a diagnostic analysis of AI scoring behavior. Results indicate that human raters exhibit significantly superior evaluative precision, with average discrimination estimates approximately double those of the AI models. Furthermore, LLMs are prone to pronounced centrality effects and score compression, systematically failing to award the highest rubric tiers. These findings demonstrate that low human-machine agreement stems from both a deficit in discriminative accuracy and systematic shifts in response criteria. Ultimately, this research provides a robust framework for calibrating and selecting AI scoring systems based on specific pedagogical goals and fairness requirements.

Xiaoliang Zhou · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.