Skip to content
Open access

Comparative performance of ChatGPT-5.2 and Gemini 3 Pro in orthopedics questions of the medical specialization examination

Jul 2026 · Anatolian Current Medical Journal · 0 citations · 23 references

Abstract

Aims: Large language models (LLMs) have demonstrated promising performance in medical examinations; however, their effectiveness in orthopedics-particularly in visually demanding questions-remains unclear. This study aimed to compare the performance of ChatGPT-5.2 and Gemini 3 Pro on orthopedics and traumatology questions from the Turkish Medical Specialization Examination (TUS).Methods: A total of 7.200 TUS questions from 2006 to 2021 were screened, and 163 orthopedics-related questions with validated answer keys were included. Questions were categorized into seven subspecialties and classified as text-based (n=154) or visual (n=9). Each question was independently presented to both models in Turkish without prior prompting. Responses were recorded as correct or incorrect. Statistical analyses included the McNemar test for paired comparisons and Fisher’s exact test for within-model comparisons of visual versus text-based performance.Results: ChatGPT-5.2 achieved an accuracy of 93.3% (152/163), while Gemini 3 Pro achieved 92.0% (150/163), with no statistically significant difference between models (p=0.593). Both models showed higher error rates on visual questions than on text-based questions (ChatGPT-5.2: 44.4% vs. 4.5%, p=0.001; Gemini 3 Pro: 33.3% vs. 6.5%, p=0.025); however, these visual comparisons were based on only nine items and should be regarded as exploratory. Error distribution across subspecialties was similar between models, with orthopedic trauma showing the highest error frequency.Conclusion: Both ChatGPT-5.2 and Gemini 3 Pro demonstrated high overall accuracy in orthopedic examination questions, with no significant difference between models. However, both models exhibited a marked decline in performance on visually based questions, highlighting persistent limitations in multimodal reasoning. Future research should incorporate larger visual datasets to better evaluate AI performance in orthopedic contexts.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.