Skip to content
Review Open access

AI-augmented versus expert-authored multiple-choice questions: a psychometric comparison in a high-stakes specialty examination

Aug 2026 · BMC Medical Education · 0 citations

Abstract

Multiple-choice questions (MCQs) are widely used in written assessments, particularly in high-stakes medical examinations. Developing high-quality MCQs is time-consuming and requires subject matter expertise. Large language models (LLMs), such as ChatGPT-4o, have therefore been proposed as tools to support item generation. Prior studies have examined AI-generated MCQs in formative or simulated contexts; evidence from authentic high-stakes examinations remains limited. This study compared the psychometric performance of AI-augmented and expert-authored MCQs within a national specialty examination. For a national high-stakes examination, 90 AI-augmented items (generated with ChatGPT-4o) and 25 expert-authored items were developed and underwent expert screening, revision, and blinded review board evaluation. The examination was administered to 304 candidates (166 basic-level, 138 specialist-level). Item analysis included 12 AI-augmented items, 14 expert-authored items, and 89 previously used bank items. Psychometric performance was assessed using classical test theory (item difficulty, discrimination, internal consistency, distractor functioning). Differences across item sources and candidate groups were analyzed using ANOVA with post-hoc testing. AI-augmented items were easier than both newly expert-authored items and previously used bank items, as reflected by a significantly higher proportion of correct responses (mean P-value: 84,7% vs. 73,5% and 69,4%, respectively; (F (2,604) = 297.187, p<0.001)). Discrimination indices were comparable across item sources. Distractor analysis showed a higher proportion of AI-augmented items with no functioning distractor ( p  < 0.05). Overall test reliability did not significantly change after inclusion of AI-augmented items. In this study, AI-augmented MCQs could be integrated into a high-stakes examination following rigorous multi-stage review, without evidence of detrimental effects on overall test performance. However, even after expert revision, such items tended to be easier and showed weaker distractor functioning. The substantial exclusion rate and systematic differences in item difficulty highlight the need for structured quality assurance and difficulty calibration when integrating AI-augmented item development into specialty-level assessments.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.