ChatGPT-4o emerged as the most reliable model for andrology-related questions, providing responses that are were highly consistent with current clinical guidelines, suggesting that error-awareness mechanisms in large language models have the potential for further refinement.
Abstract
Objectives
Large language models (LLMs) are increasingly being used in medical research and as clinical decision-support tools. This study aimed to compare the accuracy and reliability of responses generated by large language models in response to andrology-related questions.
Materials And Methods
Seventy questions concerning diagnosis, treatment, and general information were developed on the basis of the 2024 Andrology Guidelines of the European Association of Urology (EAU). These questions were submitted to three large language models, namely ChatGPT-4, ChatGPT-4o, and Google Gemini 2.0 Flash. The responses were independently evaluated by three senior urologists using a four-point rating scale. A total score (TS) > 9 indicated a good response, 6 ≤ TS ≤ 9 indicated a moderate response, and TS < 6 indicated a poor response. In addition, the self-correction capabilities of the models were evaluated, and changes in response accuracy after re-evaluation were analyzed.
Results
ChatGPT-4o achieved the highest total scores in the diagnosis and treatment categories (p < 0.001). Google Gemini 2.0 Flash generated the longest responses but demonstrated the lowest accuracy. ChatGPT-4o also showed the greatest improvement following the self-correction process (Cohen's d = - 1.214, p < 0.01). Fleiss' kappa coefficient values ranged from 0.61 to 0.80, indicating substantial interrater agreement among the urologists.
Conclusion
ChatGPT-4o emerged as the most reliable model for andrology-related questions, providing responses that are were highly consistent with current clinical guidelines. The self-correction capabilities of the models improved response accuracy, suggesting that error-awareness mechanisms in large language models have the potential for further refinement. Nevertheless, expert supervision remains essential for the safe implementation of AI-assisted systems in clinical practice.
Final recommendation concordance did not meet the protocol-defined benchmark, ChatGPT-4o never altered an MTB decision, and clinically relevant errors occurred even among highly concordant outputs, showing that concordance alone does not guarantee safety.
J. De la Torre-Trillo, Albert Munuera, M. D. Ureña et al.· Clinical and Translational O...· 0 citations
Although the four LLMs generally provide accurate and pertinent information regarding type 2 diabetes, enduring limits in actionability and inconsistencies among models in content completeness and understandability restrict their effective use in diabetic patient education.
Zhaoxia Huang, Yuxin Bai, Jun-Yue Luo et al.· Frontiers in Public Health· 0 citations
OBJECTIVE
To evaluate the quality of responses from four publicly available LLMs (ChatGPT-4o, Claude 3.7, Gemini 2.5, and Copilot) to frequently asked questions (FAQs) in pediatric urology.
METHODS
FAQs were generated using standardized prompts and submitted to each LLM using parent-centered instructions. Two board-c...
Najva Mazhari, Andrew Freedman, Nadine A. Friedrich et al.· Urology· 0 citations
Background Breast cancer, the most common malignancy among women, remains a major global public health concern. With the rapid growth of artificial intelligence–based language models, it is essential to evaluate their potential roles in patient education. This study compared ChatGPT and Gemini in responding to patient-...
Yeliz Yılmaz Bozok, N. Acar, M. Atahan et al.· PLoS ONE· 0 citations
The EAU Guidelines Bot showed excellent accuracy, completeness, and clarity when applied to neuro-urology guideline-based questions, and was comparable to that of ChatGPT 5.5, with both systems providing highly accurate guideline-concordant responses.
S. De Cillis, Riccardo Lombardo, D. Amparore et al.· European Urology Focus· 0 citations
Artificial intelligence (AI) has advanced rapidly in healthcare, with large language models (LLMs) like ChatGPT-4o showing potential in education and clinical support. This study evaluated the performance of ChatGPT-4o on oral surgery-related questions from the Japanese National Dental Examination, focusing on how visu...
H. Fukuda, M. Morishita, O. Takahashi et al.· International Journal of Ora...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.