Validity of Large Language Model Comparative Judgment for Universal Writing Screening
Universal writing screening requires scoring approaches that are feasible and supported by validity evidence. This study evaluated large language model (LLM)-based comparative judgment (CJ) for scoring informational writing assessments completed by 1,208 students in Grades 3–6 across three screening occasions. Seven LLMs representing different capability and cost tiers performed pairwise comparisons of writing quality. LLM-based CJ scores showed meaningful convergence with researcher analytic rubric scores (r = .59–.73). For single-wave scores, LLM-based CJ was comparable to researcher scoring in predicting state writing rubric scores and generally matched or exceeded it in predicting ELA scale scores and classifying ELA proficiency. Averaging scores across three screening waves substantially improved criterion-related validity and classification accuracy, with LLM-based CJ reaching β = .59–.66 for state writing rubric scores, β = .68–.74 for ELA scale scores, and AUC = .82–.86 for ELA proficiency. Predictive bias patterns for multilingual learners were similar across scoring methods. Findings were broadly consistent across LLMs, with little evidence that greater model capability or cost improved validity evidence. Results support LLM-based CJ as a promising approach for efficient writing screening and highlight the value of multiple writing samples per student and task-specific evaluation of validity relative to cost.