Skip to content

Author

Birsel Molu

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

Do AI chatbots provide reliable information on Pediatric first aid? A comparative evaluation of large language models.

PURPOSE Childhood accidents are among the leading causes of injury during early childhood. This study aimed to evaluate and compare the accuracy, clarity, and comprehensiveness of pediatric first aid information generated by LLMs. METHODS A cross-sectional comparative evaluation design was employed. Twenty standardized pediatric first aid questions were developed based on international guidelines and expert consensus. Responses generated by ChatGPT, Claude, Gemini, and Copilot were independently evaluated by a pediatric nurse and a physician using a 5-point Likert scale. Inter-rater reliability was assessed using Cohen's kappa coefficient, and differences among models were analyzed using one-way analysis of variance. RESULTS Moderate inter-rater agreement was observed across all evaluation domains. Statistically significant differences were identified among the four LLMs. Claude demonstrated the highest overall performance across all evaluation domains. Gemini demonstrated relatively high accuracy but lower clarity and comprehensiveness scores. Copilot performed well in clarity but showed limited depth of clinical content. ChatGPT received the lowest scores across all assessed domains. CONCLUSIONS The findings reveal considerable variability in the quality of pediatric first aid information generated by LLMs. While certain models may serve as supportive educational tools, none should be considered a substitute for professional medical assessment or emergency care.

Birsel Molu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.