Skip to content
Open access

Evaluation of AI Chatbot Responses to Pediatric Urology Frequently Asked Questions.

Aug 2026 · Urology · 0 citations · 15 references
Medicine

Abstract

Objective

To evaluate the quality of responses from four publicly available LLMs (ChatGPT-4o, Claude 3.7, Gemini 2.5, and Copilot) to frequently asked questions (FAQs) in pediatric urology.

Methods

FAQs were generated using standardized prompts and submitted to each LLM using parent-centered instructions. Two board-certified pediatric urologists independently rated responses across seven domains: Accuracy, Completeness, Safety, Clarity, Actionability, Conciseness, and Global Quality Score using a 5-point Likert scale. Emotional tone was analyzed using the NRC Word-Emotion Association Lexicon, and readability was assessed using seven established metrics.

Results

All LLMs produced generally high expert ratings for accuracy (mean 4.27), safety (4.36), and clarity (4.42). Claude achieved the highest overall quality score (4.35), followed by Gemini (4.03), ChatGPT (4.00), and Copilot (3.80). Emotional tone was predominantly positive and supportive across models. Reading corresponded to an 11th-15th grade reading level, exceeding the recommended 6th-8th grade patient-education standards.

Conclusion

LLMs can provide accurate, safe, and supportive information for parents, but their usefulness is limited by gaps in completeness and high readability levels. Claude achieved the highest overall quality, whereas ChatGPT demonstrated high safety with lower completeness. Future development would prioritize plain-language optimization, context-aware emotional framing, and parent co-design to improve comprehension.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.