Skip to content
Review Open access

Agreement, calibration, and failure of three large language models as high-stakes multimodal ospe graders: a comparative psychometric analysis

Jul 2026 · Medical Education Online · Vol 31 · 1 citation · 70 references
Medicine

TL;DR

Strong rank-order correlation alone does not support LLM grading for categorical decisions in high-stakes assessment contexts, and item-level gap analysis highlighted three recurring divergence patterns: non-standard histological staining, three-dimensional spatial reasoning from two-dimensional images, and semantic inflexibility in short-answer evaluation.

Abstract

Abstract Background Medical education faces increasing demand for scalable grading solutions. Large language models have been proposed as automated graders for high-stakes assessments, but evidence for their reliability in multimodal practical examinations remains limited. Methods This retrospective inter-rater reliability study compared three LLMs: ChatGPT-4o, Gemini 2.5 Flash, and Claude 3.5 Haiku with original human reference scores on 19 integrated anatomy, histology, and physiology objective structured practical examination items completed by 309 pre-medical students. Human reference scores were assigned during live summative grading by two content experts who divided items, with each item marked by one expert across all students; no human inter-rater reliability estimate was available. Items required simultaneous image interpretation and short-answer responses. Models were evaluated using zero-shot prompting reflecting deployment-realistic conditions. Agreement was assessed using Spearman's ρ, Cohen's κ, intraclass correlation coefficients, and Bland-Altman analysis. Item-level gap analysis compared student success rates with LLM performance across items. Results Rank-order correlations were strong across all models (ρ = 0.784–0.921), but categorical agreement diverged substantially. Pass/fail agreement ranged from 52.1% (ChatGPT; κ = 0.067, slight) to 84.1% (Claude; κ = 0.680, substantial). Bland-Altman analysis showed inconsistent systematic bias: Claude over-scored by +2.5 points, Gemini under-scored by −5.0 points, and ChatGPT by −10.0 points. Item-level gap analysis highlighted three recurring divergence patterns: non-standard histological staining, three-dimensional spatial reasoning from two-dimensional images, and semantic inflexibility in short-answer evaluation. The most extreme case was a pelvic three-dimensional model item on which 93.3% of students succeeded by human grading but all three LLMs assigned a mean score of zero. Conclusions Strong rank-order correlation alone does not support LLM grading for categorical decisions in high-stakes assessment contexts. LLMs may serve as assessment assistants for first-pass ranking and discrepancy flagging, but item-level human review of flagged discordances, error-profile monitoring, and preserved human authority over pass/fail decisions are required before high-stakes deployment.

Read PDF

Similar papers

Review Open access Jul 2026

Evaluating the reliability, quality, and readability of AI-generated patient education on hallux valgus: a comparative study of large language models

Patients with hallux valgus increasingly seek health information through consumer-facing artificial intelligence (AI)–driven patient education tools, particularly large language model–based conversational agents. Although these tools offer rapid and accessible responses, concerns remain regarding the reliability, usefulness, overall quality, and readability of AI-generated patient education materials. Evidence specifically evaluating AI-generated patient education for hallux valgus, a condition strongly influenced by patient expectations and treatment preferences, remains limited. This cross-sectional comparative study evaluated the performance of three large language models—ChatGPT-4o, Gemini-2.5-Flash, and DeepSeek-V3—in responding to 20 patient-centered questions related to hallux valgus. Questions were developed using AI-assisted question generation and publicly available Google Trends search patterns and categorized into four clinical domains. AI-generated responses were anonymized and independently assessed by three orthopaedic surgeons for reliability, usefulness, and overall quality using 7-point Likert-based reliability and usefulness scales and the Global Quality Scale (GQS). Readability was analyzed using six standardized indices. Inter-rater agreement and between-model comparisons were statistically evaluated. Gemini-2.5-Flash demonstrated modestly higher overall reliability, particularly in questions related to etiology and clinical presentation. DeepSeek-V3 achieved higher usefulness scores in the long-term outcomes and quality-of-life domain and produced significantly more readable content, as reflected by higher Flesch Reading Ease scores and lower grade-level indices. In contrast, Gemini-2.5-Flash generated linguistically more complex responses requiring higher educational levels for comprehension. Overall usefulness and global quality scores did not differ significantly among models. Qualitative review also identified occasional examples of oversimplified or potentially misleading information. Despite these differences, several readability metrics exceeded recommended patient health-literacy thresholds. Contemporary AI-based conversational agents can provide patient-oriented information regarding hallux valgus with variable reliability and readability characteristics, although statistically significant differences were observed across models. A trade-off between factual accuracy and linguistic accessibility was observed. AI tools should therefore be regarded as adjuncts to, rather than replacements for, clinician-led patient education. Awareness of AI limitations and appropriate clinical guidance remain essential to ensure safe, accurate, and patient-centered information delivery. Not applicable.

A. Koluman, Ebru Aloğlu Çiftçi, Mehmet Utku Çiftçi et al. · 0 citations
Open access Aug 2026

How consistent is the algorithm? Examining the intra- and inter-rater reliability of LLM-based writing assessment

Unlike constrained constructed-response assessments, the reliability of essay-based assessments has long been debated in language education research. Recently, AI—particularly large language models—has been proposed as a tool to enhance scoring reliability. This study examines the reliability of LLM-based scoring in writing assessment. OpenAI's ChatGPT assessed 192 essays written by EFL learners using a criterion-based analytic rubric covering task achievement, grammatical range and accuracy, lexical resources, organization, and mechanics. Scoring was replicated after three weeks under similar conditions. The two rounds were compared to assess intra-rater reliability and then contrasted with human ratings to assess inter-rater reliability. Results indicated high intra-rater reliability, particularly for task achievement and organization. Lower consistency was observed for mechanics and grammatical accuracy, suggesting a need for clearer criteria or further model training. Human–AI comparisons showed slight to fair exact agreement (Cohen's kappa) and moderate agreement when accounting for the degree of disagreement.

Bayan AlAmir · 0 citations
Open access Aug 2026

A benchmark dataset with human validation for AI-assisted technical answer evaluation

Grading open-ended technical responses has been a longstanding issue in higher education. Despite advances in automated assessment, existing approaches often rely on holistic scoring, weakly validated annotations, and cognitive alignment, limiting pedagogical reliability and classroom adoption. To facilitate reliable and rubric-based automated evaluation in the Data Structures and Algorithms course, this study presents DSA-RubricEval, a pedagogically grounded and preliminarily reliability-assessed dataset. The dataset constitutes the primary contribution of this work, providing a structured benchmark aligned with Bloom’s taxonomy and validated through multi-rater Interclass Correlation Coefficient ICC analysis. The dataset consists of twelve expert-designed, Bloom-aligned, questions scored on multiple rubric dimensions using an ordinal scale. Five independent evaluators scored student responses, enabling rigorous validation of human judgement based on the Intraclass Correlation Coefficient (ICC) analysis. The results indicate good to excellent average-measure reliability across most rubric dimensions, justifying the use of aggregated human scores as aggregated reference labels for automated assessment. Automated scoring was explored as a proof of concept to demonstrate the applicability of the dataset for AI-assisted assessment and formulated as an ordinal, rubric-level prediction task and evaluated using pedagogically motivated agreement measures, showing high tolerance-based agreement with human judges. The proposed research identifies sources of assessor subjectivity and explores methods to mitigate them, while reducing the workload of grading as well as correlate with learning outcomes. Moreover, the proposed technique supports lower-order cognitive skills but is not best suited for higher-order cognitive tasks. Overall, this study introduces a preliminarily reliability-assessed, rubric-based benchmark dataset intended to support exploratory research on pedagogically meaningful AI-assisted assessment of open-ended technical responses.

J. Sheikh, Hemant Kumar Soni · 0 citations
Open access Mar 2026

Evaluating Large Language Model-Based Automated Scoring in a Voice-Based Virtual Standardized Patient Platform for Medical Students: A Cross-Sectional Agreement Study.

LLM-based scoring in a VSP showed moderate agreement with faculty ratings, performing better for information gathering than for communication, and is suitable for formative use and enhanced sampling in programmatic assessment, but not for independent high-stakes summative decisions.

Xiao-Xing Gao, Xiaoming Huang, R. Hu et al. · 0 citations
Open access Aug 2026

Are large language models such as ChatGPT, capable of supporting patients and general practitioners after spine surgery?

LLMs can support communication and education following spine surgery when used with structured prompting when used with structured prompting and ChatGPT and Claude showed the highest correctness and completeness, particularly for practitioner-directed answers.

S. Wegmann, T. Rosenkranz, Philipp Egenolf et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.