Skip to content
Open access

The doctors of the future: the competition of ChatGPT-4, ChatGPT-4 omni, and Gemini 2.0 Flash in andrology.

Jul 2026 · BMC Urology · 0 citations
Medicine

TL;DR

ChatGPT-4o emerged as the most reliable model for andrology-related questions, providing responses that are were highly consistent with current clinical guidelines, suggesting that error-awareness mechanisms in large language models have the potential for further refinement.

Abstract

Objectives

Large language models (LLMs) are increasingly being used in medical research and as clinical decision-support tools. This study aimed to compare the accuracy and reliability of responses generated by large language models in response to andrology-related questions.

Materials And Methods

Seventy questions concerning diagnosis, treatment, and general information were developed on the basis of the 2024 Andrology Guidelines of the European Association of Urology (EAU). These questions were submitted to three large language models, namely ChatGPT-4, ChatGPT-4o, and Google Gemini 2.0 Flash. The responses were independently evaluated by three senior urologists using a four-point rating scale. A total score (TS) > 9 indicated a good response, 6 ≤ TS ≤ 9 indicated a moderate response, and TS < 6 indicated a poor response. In addition, the self-correction capabilities of the models were evaluated, and changes in response accuracy after re-evaluation were analyzed.

Results

ChatGPT-4o achieved the highest total scores in the diagnosis and treatment categories (p < 0.001). Google Gemini 2.0 Flash generated the longest responses but demonstrated the lowest accuracy. ChatGPT-4o also showed the greatest improvement following the self-correction process (Cohen's d = - 1.214, p < 0.01). Fleiss' kappa coefficient values ranged from 0.61 to 0.80, indicating substantial interrater agreement among the urologists.

Conclusion

ChatGPT-4o emerged as the most reliable model for andrology-related questions, providing responses that are were highly consistent with current clinical guidelines. The self-correction capabilities of the models improved response accuracy, suggesting that error-awareness mechanisms in large language models have the potential for further refinement. Nevertheless, expert supervision remains essential for the safe implementation of AI-assisted systems in clinical practice.

Read PDF

Similar papers

Review Aug 2026

ChatGPT-4o as a decision-support tool in a urological tumour board: a prospective evaluation.

Final recommendation concordance did not meet the protocol-defined benchmark, ChatGPT-4o never altered an MTB decision, and clinically relevant errors occurred even among highly concordant outputs, showing that concordance alone does not guarantee safety.

J. De la Torre-Trillo, Albert Munuera, M. D. Ureña et al. · 0 citations
Open access Aug 2026

Evaluating large language models using the Type 2 Diabetes Health Education guideline: a comparative analysis of ChatGPT-4.1, Claude-4.0, DeepSeek-V3, and ERNIE Bot 4.5 Turbo

Although the four LLMs generally provide accurate and pertinent information regarding type 2 diabetes, enduring limits in actionability and inconsistencies among models in content completeness and understandability restrict their effective use in diabetic patient education.

Zhaoxia Huang, Yuxin Bai, Jun-Yue Luo et al. · 0 citations
Open access Aug 2026

Evaluation of AI Chatbot Responses to Pediatric Urology Frequently Asked Questions.

OBJECTIVE To evaluate the quality of responses from four publicly available LLMs (ChatGPT-4o, Claude 3.7, Gemini 2.5, and Copilot) to frequently asked questions (FAQs) in pediatric urology. METHODS FAQs were generated using standardized prompts and submitted to each LLM using parent-centered instructions. Two board-c...

Najva Mazhari, Andrew Freedman, Nadine A. Friedrich et al. · 0 citations
Open access Sep 2026

Comparative evaluation of ChatGPT and gemini responses to patient-oriented questions on breast cancer

Background Breast cancer, the most common malignancy among women, remains a major global public health concern. With the rapid growth of artificial intelligence–based language models, it is essential to evaluate their potential roles in patient education. This study compared ChatGPT and Gemini in responding to patient-...

Yeliz Yılmaz Bozok, N. Acar, M. Atahan et al. · 0 citations
Review Aug 2026

Accuracy, Completeness, and Clarity of an AI-Based Chatbot for the EAU Neuro-Urology Guidelines.

The EAU Guidelines Bot showed excellent accuracy, completeness, and clarity when applied to neuro-urology guideline-based questions, and was comparable to that of ChatGPT 5.5, with both systems providing highly accurate guideline-concordant responses.

S. De Cillis, Riccardo Lombardo, D. Amparore et al. · 0 citations
Case report Open access Sep 2026

Evaluation of the performance of ChatGPT-4o on oral surgery-related questions in the Japanese National Dental Examination.

Artificial intelligence (AI) has advanced rapidly in healthcare, with large language models (LLMs) like ChatGPT-4o showing potential in education and clinical support. This study evaluated the performance of ChatGPT-4o on oral surgery-related questions from the Japanese National Dental Examination, focusing on how visu...

H. Fukuda, M. Morishita, O. Takahashi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.