Skip to content
Open access

Safety, accuracy, empathic communication, information quality, and readability of five large language model interfaces answering public questions about interstitial cystitis/bladder pain syndrome

Aug 2026 · Frontiers in Public Health · Vol 14 · 0 citations · 31 references
Medicine

TL;DR

Publicly accessible LLM interfaces showed domain-specific differences when answering IC/BPS-related public questions and may support general IC/BPS education and question preparation but should not replace clinician-led evaluation or individualized medical advice.

Abstract

Background/objectives Patients increasingly use large language model (LLM) interfaces for health information, but their safety and quality for public questions about interstitial cystitis/bladder pain syndrome (IC/BPS) remain uncertain. This study evaluated the safety, accuracy, empathic communication, information quality, reliability, and readability of five publicly accessible LLM interfaces. Methods In this CHART-guided cross-sectional comparative study, 58 public-facing IC/BPS questions were submitted once, in English, to ChatGPT, Gemini, Microsoft Copilot, DeepSeek, and Doubao using a standardized single-turn, zero-shot protocol. Three blinded senior urologists independently assessed safety, accuracy, empathic communication, DISCERN, EQIP, JAMA benchmark criteria, and Global Quality Score. Six readability indices were calculated. Results Inter-rater agreement was significant for all manually assessed metrics. Fleiss’ kappa for safety was 0.822, and ICC(2,1) values for other rater-assessed metrics ranged from 0.761 to 0.848. Unsafe responses occurred in all interfaces, ranging from 5.2% for ChatGPT to 8.6% for DeepSeek and Doubao, without a significant between-interface difference (Cochran Q = 1.000, p = 0.910). Accuracy and empathic communication differed significantly across interfaces (both p < 0.001). ChatGPT had the highest median accuracy score, whereas DeepSeek had the highest empathic communication score. Information-quality and reliability scores also differed significantly (all p < 0.001); ChatGPT achieved higher DISCERN, EQIP, and GQS scores, while Gemini achieved higher JAMA scores. Readability differed significantly across interfaces, but none met predefined patient-facing readability benchmarks. Conclusion Publicly accessible LLM interfaces showed domain-specific differences when answering IC/BPS-related public questions. Unsafe responses were uncommon but present in all interfaces, and no interface consistently outperformed the others. LLM interfaces may support general IC/BPS education and question preparation but should not replace clinician-led evaluation or individualized medical advice.

Read PDF

Similar papers

Open access Aug 2026

Safety and quality of public chatbots for lung cancer prognostic information: a comparative evaluation

Public-facing chatbots may support general patient education but should not replace individualized clinician-led prognostic communication as public-facing chatbots differed substantially in safety, reliability, communication quality, and readability.

Yan-Ru Jiang, Qian-Yun Wang, Liang Zheng et al. · 0 citations
#small language model Open access Sep 2026

Safety, accuracy, empathy, information quality, and readability of publicly accessible LLM-based chatbots for traumatic brain injury and concussion questions: a cross-sectional comparative study

Background Large language model (LLM)-based chatbots are increasingly used by the public to obtain health information, but their performance in answering questions related to traumatic brain injury (TBI) and concussion remains unclear. This study evaluated five publicly accessible LLM-based chatbots across safety, accu...

Xin Zuo, Huan Zuo, Min Zhang et al. · 0 citations
Sep 2026

Evaluation of the accuracy and reproducibility of large language models (ChatGPT, DeepSeek, Gemini) in responding to patient-centered lipedema questions.

BackgroundLipedema is a frequently misdiagnosed chronic condition that significantly impacts patients' quality of life. As artificial intelligence (AI)-based large language models (LLMs) become increasingly integrated into healthcare communication, their accuracy and consistency in providing patient-centered informatio...

Rabia Sanır, E. Türkmen, E. Giray et al. · 0 citations
Open access Aug 2026

Quality, readability, and clinical-risk signals of default public-interface LLM responses to vascular and perioperative patient questions: a cross-sectional snapshot

In this English-language benchmark of default first responses from five public LLM interfaces accessed through specific logged-in accounts from a Hong Kong, China IP address during a single late-May 2026 window, 107 of 110 responses received the maximum Guideline Concordance Score, and no response met the prespecified...

Wei Zhong, Yuanyuan Zhang, Yu Huang et al. · 0 citations
Review Open access Aug 2026

Quality, readability, and patient safety of ChatGPT-generated responses to fall-related questions in older adults: a multidisciplinary evaluation

While ChatGPT provided generally acceptable clinical information, variability in readability and expert ratings raises patient safety concerns and AI-generated health content should be reviewed and tailored to older adults' health literacy needs before clinical use.

Merve Arı, N. Ilçin, Hatice Yağcıoğlu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.