Evaluating Large Language Model-Based Automated Scoring in a Voice-Based Virtual Standardized Patient Platform for Medical Students: A Cross-Sectional Agreement Study.
Mar 2026· JMIR Medical Education· 0 citations· 22 references
Medicine
TL;DR
LLM-based scoring in a VSP showed moderate agreement with faculty ratings, performing better for information gathering than for communication, and is suitable for formative use and enhanced sampling in programmatic assessment, but not for independent high-stakes summative decisions.
Abstract
Background
Large language model (LLM)-powered virtual standardized patients (VSPs) enable scalable clinical skills practice, but the validity of AI-generated scores relative to faculty ratings remains unclear.
Objective
To assess agreement between LLM-generated and faculty ratings of history-taking and communication performance, and to examine the influence of rater and case heterogeneity.
Methods
In this cross-sectional study, 92 fourth-year medical students completed one of three 15-minute voice-based VSP cases (fever, diarrhea, cough). Ten blinded faculty raters scored performance (0-100 total; 0-50 domains). AI scores were generated by DeepSeek-V3 using a calibrated prompt. Agreement was evaluated using mixed-effects models, intraclass correlation coefficients (ICC[2,1]), Spearman correlations, mean absolute error (MAE), Bland-Altman analysis, and variance partition coefficients (VPC).
Results
Median total scores were similar for AI and faculty (93.0 [IQR 6.0] vs 94.0 [IQR 4.0]). Rater variability accounted for 37% of residual variance in faculty total scores (VPC = 0.37). AI total scores were positively associated with faculty total scores (β = 0.37, 95% CI 0.26-0.48, P<.001; Spearman ρ = 0.50, 95% CI 0.34-0.65). Absolute agreement was moderate (ICC[2,1] = 0.51, 95% CI 0.34-0.65), with MAE of 3.11 points. Mixed-effects Bland-Altman analysis showed a non-significant mean bias (1.26 points) and 95% limits of agreement from -4.95 to 7.48 (width = 12.43 points), with proportional bias (β_mean = -0.55, P<.001). Agreement was stronger for information gathering (β = 0.46, ρ = 0.49, ICC = 0.54, VPC = 0.23) than for communication (β = 0.27, ρ = 0.28, ICC = 0.29, VPC = 0.52). A sensitivity analysis in the lowest quartile showed attenuated but consistent agreement (ICC = 0.38).
Conclusions
LLM-based scoring in a VSP showed moderate agreement with faculty ratings, performing better for information gathering than for communication. Due to rater and case heterogeneity, ceiling effects, and proportional bias, it is suitable for formative use and enhanced sampling in programmatic assessment, but not for independent high-stakes summative decisions.
CLINICALTRIAL
BACKGROUND
Medical history taking (MHT) is a foundational clinical competency for medical students; however, traditional training models using standardized patients face challenges such as resource constraints. Large language model-powered virtual standardized patients (LLM-VSPs) offer a safe, repeatable platform for self-directed practice with AI-automated feedback. Nevertheless, their effectiveness in authentic teaching environments and underlying learning mechanisms require further investigation.
OBJECTIVE
This study aims to evaluate the impact of an LLM-VSP system as an extracurricular self-practice tool on undergraduate medical students' MHT performance in an authentic educational setting without disrupting routine instruction, and to further explore potential associations among practice behaviors, baseline proficiency, and intervention effects.
METHODS
This prospective cohort study enrolled 168 third-year medical students. Based on voluntary participation, students were assigned to an intervention group (n=120, using LLM-VSP) or a control group (n=48, receiving routine instruction). Propensity score matching (PSM) balanced confounding factors, yielding 40 matched pairs. Baseline MHT performance was assessed via virtual patient examination after didactic instruction but before clinical practicum. The primary outcome was end-of-term MHT performance assessed at an Objective Structured Clinical Examination station with real standardized patients. The differences between groups were compared using independent samples t test, with robustness validated through multiple linear regression, sensitivity analyses, and Rosenbaum bounds analyses. Exploratory analyses investigated the association between practice behaviors and scores, and observed benefit differences across baseline levels.
RESULTS
After PSM, baseline characteristics were balanced (standardized mean difference <0.1). The intervention group exhibited higher total MHT scores than controls (mean 87.71, SD 7.29 vs mean 83.74, SD 8.06; mean difference 3.98, 95% CI 0.55-7.40 points; P=.02; with a medium effect size of Cohen d=0.52). Advantages were observed in content (P=.03) and communication skills (P=.04) subscores. Regression analysis confirmed robust intervention effects (B=3.924, 95% CI 1.90-5.94; P<.001; R2=0.708), with sensitivity analyses supporting reliability. Exploratory analysis suggested that mere practice behavior metrics were not independent predictors of final scores, potentially being constrained by baseline proficiency, implying a "cognitive threshold" for effective AI-assisted training. Subgroup analysis indicated a trend of differential benefits: high baseline students appeared to show larger gains (matched sample: mean difference 5.93, 95% CI 2.17-9.70; overall sample: mean difference 7.04, 95% CI 3.84-10.25), whereas improvements in medium and low baseline students were relatively limited (matched sample: mean difference 2.94-2.99; overall sample: mean difference 1.29-1.64).
CONCLUSIONS
Introducing LLM-VSPs as a self-practice tool in diagnostics education may help improve undergraduate medical students' MHT performance. Preliminary evidence suggests a potential "cognitive threshold," implying students with solid theoretical foundations and higher baseline proficiency may better achieve skill transformation through AI-assisted autonomous practice. This provides a basis for future stratified teaching strategies and differentiated guidance for students of varying baseline levels.
Unknown authors· JMIR Medical Education· 0 citations
Patients with hallux valgus increasingly seek health information through consumer-facing artificial intelligence (AI)–driven patient education tools, particularly large language model–based conversational agents. Although these tools offer rapid and accessible responses, concerns remain regarding the reliability, usefulness, overall quality, and readability of AI-generated patient education materials. Evidence specifically evaluating AI-generated patient education for hallux valgus, a condition strongly influenced by patient expectations and treatment preferences, remains limited.
This cross-sectional comparative study evaluated the performance of three large language models—ChatGPT-4o, Gemini-2.5-Flash, and DeepSeek-V3—in responding to 20 patient-centered questions related to hallux valgus. Questions were developed using AI-assisted question generation and publicly available Google Trends search patterns and categorized into four clinical domains. AI-generated responses were anonymized and independently assessed by three orthopaedic surgeons for reliability, usefulness, and overall quality using 7-point Likert-based reliability and usefulness scales and the Global Quality Scale (GQS). Readability was analyzed using six standardized indices. Inter-rater agreement and between-model comparisons were statistically evaluated.
Gemini-2.5-Flash demonstrated modestly higher overall reliability, particularly in questions related to etiology and clinical presentation. DeepSeek-V3 achieved higher usefulness scores in the long-term outcomes and quality-of-life domain and produced significantly more readable content, as reflected by higher Flesch Reading Ease scores and lower grade-level indices. In contrast, Gemini-2.5-Flash generated linguistically more complex responses requiring higher educational levels for comprehension. Overall usefulness and global quality scores did not differ significantly among models. Qualitative review also identified occasional examples of oversimplified or potentially misleading information. Despite these differences, several readability metrics exceeded recommended patient health-literacy thresholds.
Contemporary AI-based conversational agents can provide patient-oriented information regarding hallux valgus with variable reliability and readability characteristics, although statistically significant differences were observed across models. A trade-off between factual accuracy and linguistic accessibility was observed. AI tools should therefore be regarded as adjuncts to, rather than replacements for, clinician-led patient education. Awareness of AI limitations and appropriate clinical guidance remain essential to ensure safe, accurate, and patient-centered information delivery.
Not applicable.
A. Koluman, Ebru Aloğlu Çiftçi, Mehmet Utku Çiftçi et al.· BMC Medical Informatics and...· 0 citations
This exploratory scenario-based study demonstrates model-specific, rather than universal, language effects on LLM performance in dental trauma management.
This study conducts a clinical evaluation of a secure, locally deployed, quantized large language model (LLM) for automating modified Rankin Scale (mRS) score extraction from unstructured neurosurgical notes. We retrospectively selected 103 authentic clinical letters (2007–2025) from aneurysm patients at a tertiary neurosurgical centre. To comply with data privacy constraints, an open-source reasoning LLM (Qwen3-32B) with 4-bit quantization was deployed entirely on-premises. The LLM extracted mRS scores using a zero-shot approach with custom logits processors to enforce strict JSON formatting. Performance was compared to a reference standard (attending neurosurgeons’ consensus) and parallel scoring by medical residents and students. The LLM achieved excellent agreement with the attending consensus (QWK 0.95), matching the reliability of medical students (QWK 0.95) and residents (QWK 0.93). Exact agreement was 75%, and agreement within ±1 mRS point was 96%. Bayesian analysis strongly supported statistical equivalence between the model and human raters. The computationally optimized LLM demonstrated human-level classification reliability without task-specific fine-tuning. This approach successfully addresses key patient data privacy barriers and the formatting inconsistencies typical of open-ended generative models. Securely deploying a general-purpose, quantized LLM provides a scalable pathway for extracting functional outcomes and supports FAIR-aligned data systems.
A. Bornet, Abiram Sandralegar, A. Yazdani et al.· Machine Learning and Knowled...· 0 citations
Background After diagnosis of high-risk human papillomavirus (HPV) infection, patients often seek guidance on cancer risk, colposcopy, follow-up, treatment, partner management, pregnancy, and anxiety. Large language models (LLMs) are increasingly used for medical consultation, but their response quality and safety in this setting require evaluation. Methods In this single-center, expert-rated methodological study, 15 patient-style questions based on common outpatient consultations were submitted to ChatGPT-5.5 Instant and DeepSeek-V3. Each model generated three independent responses per question, yielding 90 artificial intelligence (AI)-generated responses. Five blinded gynecology experts independently evaluated all responses, producing 450 expert-rating records. Accuracy, safety, guideline concordance, completeness, and understandability were rated on a 1–5 Likert scale. Experts also assessed major error and potential harm. The primary outcome was the proportion of responses without major error or potential harm. Results ChatGPT-5.5 Instant had higher response-level scores than DeepSeek-V3 for accuracy (mean difference, 0.25; 95% CI, 0.11–0.39; FDR-adjusted p = 0.010), safety (0.27; 0.07–0.47; p = 0.010), completeness (0.22; 0.07–0.37; p = 0.020), and composite score (0.18; 0.04–0.31; p = 0.027). The difference in guideline concordance did not remain significant after multiplicity correction (0.16; −0.01 to 0.33; FDR-adjusted p = 0.057), and understandability was similar. Responses without major error or potential harm occurred in 42/45 (93.3%; 95% CI, 81.7–98.6%) ChatGPT responses and 38/45 (84.4%; 70.5–93.5%) DeepSeek-V3 responses (risk difference, 8.9 percentage points; 95% CI, −13.3 to 31.1; p = 0.344). Exploratory question-level risks were observed in HPV16/18 positivity, normal cytology, partner management, and pregnancy scenarios. Conclusion Both two models demonstrated generally strong performance across the evaluated quality domains. ChatGPT achieved higher ratings in several quality domains, but response-level binary safety differences were imprecise and not statistically significant. Both models require guideline-based clinical oversight.
Zhen Hao, Lin Wang, Yue Wu et al.· Frontiers in Public Health· 0 citations
ObjectiveThis study evaluates the performance of large language models (LLMs)-ChatGPT-4.0, Gemini 2.0 Pro, o3-mini, Doctor GPT and DeepSeek-V3-in a national orthopaedic proficiency examination and explores their implications for health informatics and medical education. The responses of these models were analysed to assess accuracy rates and differences between models.MethodA total of 100 multiple-choice questions from the 2024 TOTEK examination were administered to each AI model under identical conditions. Correct and incorrect responses were recorded, and differences in performance were evaluated using chi-square testing and frequency analysis. Question categories were also compared to identify domain-specific variations.Resultso3-mini achieved the highest accuracy rate (79%), while Gemini 2.0 showed the lowest (68%); all models exceeded the 60% pass threshold. A statistically significant difference between models was identified in the Surgical Procedures category, in which Gemini 2.0 answered fewer questions correctly (23/36) than the other models (30-32/36) (χ2 = 9.87, df = 4, p = 0.043). No significant differences were observed in the remaining categories (all p > 0.05), and the overall difference in accuracy between models did not reach statistical significance (χ2 = 4.01, df = 4, p = 0.405). Clinical decision-making and visual content-based questions were the most challenging for all models.ConclusionAI models demonstrate generally high accuracy in medical examinations; however, they struggle with interpreting clinical context, recognising atypical medical scenarios and answering questions involving visual content.
Bünyamin Arı· Health Informatics Journal· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.