Aug 2026· Advances in Medical Education and Practice· Vol 17, pp. 1-11· 0 citations· 19 references
Medicine
TL;DR
Assessing overall accuracy, test-retest reliability, topic-specific performance, and the relationship between cognitive complexity and model performance in large language models for microbiology multiple-choice questions found response consistency should be a primary criterion for educational deployment.
Abstract
Background Large language models (LLMs) are increasingly used as learning resources in medical education, yet their performance and reliability in microbiology, a discipline with a broad, heterogeneous knowledge base, have not been systematically evaluated across multiple platforms. Objective To benchmark seven publicly available LLMs on microbiology multiple-choice questions (MCQs), assessing overall accuracy, test-retest reliability, topic-specific performance, and the relationship between cognitive complexity and model performance. Methods Seven LLMs (Claude 4.6 Sonnet, Gemini 3.0, ChatGPT-5.2, Grok 4, Copilot, DeepSeek V3, and Kimi K2) completed 200 MCQs distributed across 20 microbiology topics and five Bloom’s taxonomy levels in three independent sessions separated by 24-hour intervals. A total of 4200 responses were analyzed. Statistical analysis included one-way ANOVA with Tukey’s HSD post hoc tests, repeated-measures ANOVA, intraclass correlation coefficients (ICCs), and Pearson correlations. Results The collective mean accuracy was 86.18%. Six of seven systems exceeded the 80% high-competency threshold; Claude (89.83%), Grok (89.50%), and GPT (88.83%) led the group. Gemini (72.33%) was the only underperforming system. Test-retest reliability varied dramatically: Claude achieved excellent ICC (0.966), while Gemini exhibited poor reliability (ICC = 0.290), with session-to-session fluctuations of up to 100 percentage points on individual topics. Microbial Cell (100%) was the easiest topic; Viral Genomics (61.9%) was the most challenging across all systems. A uniform decline at Bloom’s Level 4 (Analyze) was observed across all LLMs, with no model exceeding 78%. Conclusion Contemporary LLMs demonstrate substantial knowledge of microbiology but differ markedly in reliability. Response consistency, alongside accuracy, should be a primary criterion for educational deployment. These findings are specific to microbiology MCQ performance and may not generalize to open-ended clinical reasoning.
Abstract Background Large language models (LLMs) are rapidly transforming medical education, yet their performance in Allergy/Immunology remains insufficiently characterized. Furthermore, concerns regarding accuracy, consistency, and sensitivity to input format persist. Objectives This study aimed to evaluate and compare the accuracy and response consistency of three leading LLMs—ChatGPT-5, Gemini 2.5, and Grok 4—on Allergy/Immunology United States Medical Licensing Examination (USMLE) Step 1-style questions under different prompt conditions. Methods Thirty-five USMLE Step 1-style questions were selected. Questions were presented to each model in two formats: single-question prompts and a combined prompt containing all questions. Fifteen trials were conducted for each format per model. Performance was assessed using mean accuracy, and variability was measured using Shannon entropy. Mixed-effects models tested the effects of model, prompt condition, and question difficulty. Results Overall accuracy differed significantly ( p < 0.001), with Gemini (80.7%) and Grok (80.5%) achieving higher mean scores than ChatGPT (74.3%). Single-item prompts yielded superior performance, with Grok (93.1%) and Gemini (90.9%) demonstrating the highest accuracy. Transitioning to a combined prompt significantly reduced accuracy for all models. Accuracy also decreased with increasing question difficulty for all models. Grok demonstrated superior reliability, maintaining the lowest overall response entropy, whereas ChatGPT exhibited the highest variability. Conclusion On Allergy/Immunology Step 1-style questions, Gemini and Grok demonstrated higher accuracy than ChatGPT, although their overall accuracies remained approximately 81%. Grok offered the most consistent performance. All models demonstrated substantial sensitivity to prompt complexity and inherent performance limitations. These findings underscore the importance of prompt optimization and support the supplementary role of these models in medical education.
M. Carroll, Sabrina Kentis, Hannah Kareff et al.· Applied Clinical Informatics· 0 citations
LLMs show potential as tools for preliminary drug information retrieval and rapid responses generation in drug information services, however, variable concordance and persistent limitations in citation credibility indicate the need for continued pharmacist oversight.
Nuntapong Boonrit, Najwa Bin-Useng, Aphichaya Sirijariyawat et al.· Journal of the American Coll...· 0 citations
Background The increasing number of student cohorts has compelled academics to create a larger number of Multiple-Choice Question (MCQ) items. Large language models (LLMs) can help educators generate assessment items across multiple disciplines. The study compared the ability of LLMs (Gemini Advanced, Perplexity Pro, and ChatGPT 4.0) to generate high-quality, clinical-scenario-based MCQ items across three disciplines in a medical program, using an 8-point quality rubric. Materials and Methods Learning Outcomes (LOs) from Anatomy and Pathology disciplines of a pre-clinical semester 4 module and Family Medicine discipline of a clinical semester 6 module of a medical undergraduate program were selected. Using a pre-determined descriptive prompt, 63 MCQ items were generated from three AI tools. The quality of item construction was assessed by external content experts who were blinded to item generation using an 8-criterion rubric and a 4-point Likert scale. Mean scores and ranks for MCQs under each LLM were analysed, and a Friedman test was conducted to compare them. Kendall’s W showed that all criteria except one demonstrated some effect and weak-to-fair inter-rater reliability. Results When measuring key problem-solving skills, Perplexity Pro-generated MCQ items received a higher percentage of “strongly agree” ratings in Anatomy (57.1%, n=63), Pathology (56.5%, n=63), and Family Medicine (71.4%, n=63). Perplexity Pro received a higher percentage of “strongly agree” ratings in Anatomy, at 47.6% (n = 63), when evaluating specific content. Gemini Advanced also scored highly, with 68.8% in Pathology and 81% in Family Medicine (n = 63). A comparative analysis of the higher mean scores and ranks across 24 quality criteria showed that Perplexity Pro, Gemini Advanced, and ChatGPT 4.0 achieved 13, 10, and 1 higher mean scores, respectively. There is no significant difference in mean scores and ranks among the three LLMs across different quality criteria. Conclusion MCQs created by Perplexity Pro and Gemini Advanced achieved comparatively higher percentages of “strongly agree” ratings across quality criteria for item construction. Assessing LLM-generated assessment items provides valuable insight into the quality of LLM-supported MCQ questions.
N. K. Mitra, Thirupathirao. Vishnumukkala, Thin Thin Win et al.· Advances in Medical Educatio...· 0 citations
This study aims to comparatively examine the readability, accuracy, and quality of responses provided by artificial intelligence (AI)-based chatbots such as Perplexity, ChatGPT-5, and Gemini to questions about knee osteoarthritis (KOA), which accounts for approximately four-fifths of the global osteoarthritis (OA) burden. In this study, 8 keywords were determined by excluding repetitive, irrelevant or synonymous ones from the 25 most frequently used English keywords associated with KOA based on Google Trends data, and these terms were asked as questions to three different artificial intelligence-based chatbots. The study measured readability using formulas like Coleman-Liau Index (CLI), Automated Readability Index (ARI), and Linsear Write (LW). Reliability of the information was assessed using the Journal of the American Medical Association (JAMA) benchmarks along with the modified DISCERN instrument. To determine overall content quality, the Global Quality Score (GQS) and the Ensuring Quality Information for Patients (EQIP) scale were applied. Together, these tools provided a comprehensive assessment of how understandable, reliable, and high-quality each chatbot’s responses were. The most frequently searched keywords related to OA were “osteoarthritis of knee,” “knee pain,” and “osteoarthritis knee pain.” A readability analysis of responses from three different AI-based chat systems revealed that all platforms had text levels above the Grade 6 threshold, and this difference was statistically significant (p < 0.05). Comparisons demonstrated that ChatGPT-5 produced the most readable content (FRES:45, GFOG:11.9, FKGL:9.24, CLI:14.03, SMOG:8.37, ARI:11.37, LW:7.2). However, Perplexity achieved significantly higher scores than ChatGPT-5 across all quality and reliability assessments, yielding superior median scores (DISCERN: 4, JAMA: 2, GQS: 4, EQIP: 92.8). Perplexity also outperformed Gemini in the mDISCERN reliability assessment (p = 0.001), while no significant difference in quality or reliability was found between Gemini and ChatGPT-5. No statistically significant difference was found between Gemini and ChatGPT in reliability and quality surveys. This analysis of KOA highlights significant challenges regarding the potential of popular AI chatbots for patient information. When examining readability levels, responses from these tools consistently exceed the recommended comprehensibility threshold, making it difficult for patients to absorb critical information. Furthermore, the relatively low scores recorded in reliability and content quality assessments raise significant concerns about the scientific validity and integrity of the medical information presented. Given these findings, the sufficient quality, robustness, and appropriate levels of understandability of future AI-based tools can only be ensured by the establishment and operation of an effective oversight mechanism.
Erdem Maraşlı, E. Ozduran, Volkan Hancı· PLoS ONE· 1 citation
The LLM4DS-Benchmark is introduced and a multidimensional empirical evaluation of seven large language models is conducted, highlighting the need for multidimensional, task-aware benchmarking and suggesting that model selection for data science coding should be guided by task characteristics and practical constraints rather than aggregate success rate alone.
Santhosh Anitha Boominathan, Sai Sanjna Chintakunta, Everton Guimarães et al.· Empirical Software Engineeri...· 0 citations
Assessment methods are essential for evaluating medical students’ knowledge and cognitive skills. MCQs are widely used for their objectivity, while SAQs are considered better for assessing higher-order cognition, yet comparative psychometric evidence in formative settings is limited.
To compare the difficulty index (DIF), discrimination index (DI), and overall test performance of MCQs and SAQs used in a formative assessment of first-year medical students, and to examine differences in item performance across achievement levels.
Descriptive analytical cross-sectional study.
The study included 178 first-year medical students enrolled in the Medical Parasitology course within the Locomotor and Skin (MLS-1) module at Ain Shams University during the 2024–2025 academic year. Students completed both MCQ and SAQ examinations, each comprising 20 items mapped to identical intended learning outcomes. Psychometric analysis included item difficulty, discrimination indices, internal consistency reliability, and comparative statistical testing between formats.
Students scored significantly higher on MCQs than SAQs (mean 13.5 vs. 11.1 out of 20, p < .001, Cohen’s d = 0.59). MCQs were significantly easier than SAQs (mean DIF 0.67 vs. 0.55, p = .005), while discrimination indices did not differ significantly (mean DI 0.44 vs. 0.41, p = .276). Both formats demonstrated excellent reliability (Cronbach’s α > 0.80) and strong discrimination between highand low achievers. Pass/fail concordance was 66.3%, with more students passing MCQs but failing SAQs.
MCQs and SAQs show strong psychometric quality. MCQs are easier, while SAQs are more challenging yet equally discriminative, supporting their complementary use in formative medical education
Unknown authors· The Quarterly journal of med...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.