Skip to content

Author

Deborah K. Reed

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Validity of Large Language Model Comparative Judgment for Universal Writing Screening

Universal writing screening requires scoring approaches that are feasible and supported by validity evidence. This study evaluated large language model (LLM)-based comparative judgment (CJ) for scoring informational writing assessments completed by 1,208 students in Grades 3–6 across three screening occasions. Seven LLMs representing different capability and cost tiers performed pairwise comparisons of writing quality. LLM-based CJ scores showed meaningful convergence with researcher analytic rubric scores (r = .59–.73). For single-wave scores, LLM-based CJ was comparable to researcher scoring in predicting state writing rubric scores and generally matched or exceeded it in predicting ELA scale scores and classifying ELA proficiency. Averaging scores across three screening waves substantially improved criterion-related validity and classification accuracy, with LLM-based CJ reaching β = .59–.66 for state writing rubric scores, β = .68–.74 for ELA scale scores, and AUC = .82–.86 for ELA proficiency. Predictive bias patterns for multilingual learners were similar across scoring methods. Findings were broadly consistent across LLMs, with little evidence that greater model capability or cost improved validity evidence. Results support LLM-based CJ as a promising approach for efficient writing screening and highlight the value of multiple writing samples per student and task-specific evaluation of validity relative to cost.

Sterett Mercer, Deborah K. Reed · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.