Agreement, calibration, and failure of three large language models as high-stakes multimodal ospe graders: a comparative psychometric analysis
Strong rank-order correlation alone does not support LLM grading for categorical decisions in high-stakes assessment contexts, and item-level gap analysis highlighted three recurring divergence patterns: non-standard histological staining, three-dimensional spatial reasoning from two-dimensional images, and semantic inflexibility in short-answer evaluation.