Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

The temporal changes in GPT-4 performance on UKMLA practice questions: educational and clinical implications

Background Generative Artificial Intelligence (GAI) models, such as GPT-4, have been extensively studied for their integration into medical practice and education. GPT-4 has demonstrated excellent performance on medical licensing examinations, including the United Kingdom Medical Licensing Assessment (UKMLA). However, the field lacks longitudinal data on whether such performance is stable or varies over time. Theoretically, iterative improvements updated by the vendor should enhance performance, but empirical evidence of such longitudinal trends remains limited. Given that GPT-4 undergoes periodic vendor-side updates, we aimed to analyse the categorical, time-spaced performance of GPT-4 on the UKMLA to characterise how its performance changes over time in a medical context. Methods Two publicly available UKMLA papers were fed into GPT-4 at two different time points, June 2023 and December 2024. 191 questions were provided with and without multiple-choice options to assess GPT-4’s clinical competence. McNemar’s test was performed to evaluate changes in GPT-4’s performance over time, comparing domain-specific questions. Results GPT-4’s accuracy improved noticeably between the two rounds (MCQ: 88.0 to 93.7%, p = 0.027; non-MCQ: 68.1 to 81.7%, p = 1.00). Single-step accuracy rose from 73.1 to 82.3%, and multi-step from 57.4 to 80.3% without MCQ. GPT-4 showed improved accuracy from Round 1 to Round 2 for both single-step and multi-step questions, with MCQ-prompted responses consistently outperforming non-MCQ responses (up to 95.1% accuracy for multi-step MCQ questions in Round 2). GPT-4’s performance improved across all question categories from round one to round two, most notably in management questions without MCQ options (+23.30%), though these differences were not statistically significant. Discussion and conclusion GPT-4’s performance on the UKMLA improved significantly over 18 months, suggesting that iterative vendor-side model updates enhance clinical reasoning capabilities. These findings indicate that GPT-4 may serve as a supplementary educational tool for medical students and clinicians; however, the underlying drivers of performance changes remain opaque, and such tools should be deployed with structured oversight to prevent overreliance.

Ravanth Baskaran, Sai Sirikonda, Aditya Singh et al. · 0 citations
Open access Jul 2026

MerMED-FM: Multimodal, Multi-Disease Medical Imaging Foundation Model.

BACKGROUND Current artificial intelligence (AI) models for medical imaging predominantly focus on a single imaging modality and a single disease. Attempts to create multimodal and multi-disease models have resulted in inconsistent clinical accuracy. Furthermore, training these models typically requires large, well labelled datasets, which are costly and labour intensive to prepare. We aimed to train and evaluate an AI model that can interpret diverse imaging modalities across specialties while maintaining robust performance within each modality. METHODS We developed Multimodal, Multi-Disease Medical Imaging Foundation Model (MerMED-FM), a multi-specialty model trained using self-supervised learning and a memory module. MerMED-FM was pretrained on publicly sourced, unlabelled medical images from 12 specialties and seven imaging modalities: chest x-rays, CT, ultrasound, histopathology, colour fundus photography (CFP), optical coherence tomography (OCT), and dermatoscopy. After pretraining, the model was fine-tuned, validated, and evaluated for the diagnosis of a range of diseases on 26 public datasets and five private datasets comprising radiology, histopathology, and ophthalmology images. MerMED-FM was compared against a general-domain vision foundation model, various specialist single-modality foundation models, and a multispecialty foundation model. Models were fine-tuned using 10%, 30%, 50%, and 100% of data, with primary comparative analyses conducted using a 10% label fraction. The primary outcome was the area under the receiver operating characteristic curve (AUROC), which was summarised by imaging modality. FINDINGS MerMED-FM was trained on around 3·3 million images from 53 publicly available, unlabelled datasets, comprising 713 931 chest x-rays, 292 353 CT slices, 389 885 ultrasound frames, 1 017 712 pathology patches, 333 099 CFP images, 176 719 OCT slices, and 401 059 dermatoscopy images. Strong performance was achieved across all modalities at a label fraction of only 10%, with mean AUROC values of 0·844 for chest x-rays, 0·906 for CT, 0·818 for ultrasound, 0·908 for histopathology, 0·810 for CFP, 0·962 for OCT, and 0·827 for dermatoscopy. INTERPRETATION MerMED-FM has the potential to be a highly adaptable, versatile, cross-specialty foundation model that enables robust interpretation of medical imaging across diverse medical disciplines. FUNDING National Medical Research Council, Singapore and the Agency for Science, Technology and Research, Singapore.

Yang Zhou, C. Quek, Jun Zhou et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.