Open access
Mar 2026
Evaluating Large Language Model-Based Automated Scoring in a Voice-Based Virtual Standardized Patient Platform for Medical Students: A Cross-Sectional Agreement Study.
LLM-based scoring in a VSP showed moderate agreement with faculty ratings, performing better for information gathering than for communication, and is suitable for formative use and enhanced sampling in programmatic assessment, but not for independent high-stakes summative decisions.
Xiao-Xing Gao, Xiaoming Huang, R. Hu et al.
· JMIR Medical Education · 0 citations