Skip to content

Author

O. Casals-Farre

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

The temporal changes in GPT-4 performance on UKMLA practice questions: educational and clinical implications

Background Generative Artificial Intelligence (GAI) models, such as GPT-4, have been extensively studied for their integration into medical practice and education. GPT-4 has demonstrated excellent performance on medical licensing examinations, including the United Kingdom Medical Licensing Assessment (UKMLA). However, the field lacks longitudinal data on whether such performance is stable or varies over time. Theoretically, iterative improvements updated by the vendor should enhance performance, but empirical evidence of such longitudinal trends remains limited. Given that GPT-4 undergoes periodic vendor-side updates, we aimed to analyse the categorical, time-spaced performance of GPT-4 on the UKMLA to characterise how its performance changes over time in a medical context. Methods Two publicly available UKMLA papers were fed into GPT-4 at two different time points, June 2023 and December 2024. 191 questions were provided with and without multiple-choice options to assess GPT-4’s clinical competence. McNemar’s test was performed to evaluate changes in GPT-4’s performance over time, comparing domain-specific questions. Results GPT-4’s accuracy improved noticeably between the two rounds (MCQ: 88.0 to 93.7%, p = 0.027; non-MCQ: 68.1 to 81.7%, p = 1.00). Single-step accuracy rose from 73.1 to 82.3%, and multi-step from 57.4 to 80.3% without MCQ. GPT-4 showed improved accuracy from Round 1 to Round 2 for both single-step and multi-step questions, with MCQ-prompted responses consistently outperforming non-MCQ responses (up to 95.1% accuracy for multi-step MCQ questions in Round 2). GPT-4’s performance improved across all question categories from round one to round two, most notably in management questions without MCQ options (+23.30%), though these differences were not statistically significant. Discussion and conclusion GPT-4’s performance on the UKMLA improved significantly over 18 months, suggesting that iterative vendor-side model updates enhance clinical reasoning capabilities. These findings indicate that GPT-4 may serve as a supplementary educational tool for medical students and clinicians; however, the underlying drivers of performance changes remain opaque, and such tools should be deployed with structured oversight to prevent overreliance.

Ravanth Baskaran, Sai Sirikonda, Aditya Singh et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.