Skip to content
Open access

Performance evaluation of large language models in bladder cancer patient education Q&A: a cross-sectional study

Sep 2026 · Frontiers in Oncology · Vol 16 · 0 citations · 32 references
Medicine

TL;DR

Current mainstream LLMs demonstrate initial potential for generating educational content on bladder cancer, albeit with considerable heterogeneity across models, and advocate for a prudent, assistive role of LLMs in health communication under a human-AI collaborative model.

Abstract

Background Bladder cancer ranks among the most prevalent urological tumors worldwide, with its global incidence continuing to rise steadily. Although patient education materials (PEMs) play a crucial role in enhancing disease comprehension and supporting joint clinical decision-making, current online resources frequently surpass the readability thresholds recommended for the general public. Large language models (LLMs) hold promise for health communication, yet no systematic assessment has been conducted regarding their feasibility and trustworthiness specifically for bladder cancer patient education. Objective This study aimed to systematically benchmark five leading LLMs in producing question-and-answer content for bladder cancer science popularization, with a particular focus on readability, informational quality, and appropriateness for patient education. Methods In this cross-sectional simulation study, 20 common patient questions covering five disease domains were compiled. On January 15, 2026, each question was submitted identically to five publicly available LLMs (Doubao, DeepSeek, Kimi, Gemini, and ChatGPT). Readability was evaluated using seven conventional metrics. Two independent pharmacists, blinded to model identity, rated the responses using the Chinese version of the Patient Education Materials Assessment Tool for print materials (C-PEMAT-P) and the Global Quality Score (GQS). Additionally, two independent clinical specialists assessed factual accuracy and alignment with the Chinese Bladder Cancer Diagnosis and Treatment Guidelines (2024 edition) employing a 4-point scale. Cohen’s kappa was used to determine inter-rater reliability. Results ChatGPT, DeepSeek, and Doubao outperformed Kimi and Gemini on both C-PEMAT and GQS (all P < 0.001), indicating superior understandability, actionability, and overall quality. Across all models, median C-PEMAT scores ranged from 8 to 10, suggesting broadly acceptable suitability for patient education. Readability varied significantly by content domain, with treatment-oriented texts showing the highest complexity. ChatGPT achieved the best alignment with clinical guidelines. No model produced harmful advice or directly contradicted guideline recommendations. Traditional readability measures correlated weakly with GQS, whereas C-PEMAT showed a moderate positive correlation (r = 0.34). Conclusion Current mainstream LLMs demonstrate initial potential for generating educational content on bladder cancer, albeit with considerable heterogeneity across models. Disease-specific evaluation instruments for patient education materials are more effective than general readability formulas in reflecting perceived quality. Our results advocate for a prudent, assistive role of LLMs in health communication under a human-AI collaborative model.

Read PDF

Similar papers

Sep 2026

Large Language Models for Breast Cancer Education: A Comparative Analysis of Quality, Reliability and Readability.

Gemini significantly outperforms ChatGPT in response quality, reliability, and linguistic accessibility for breast cancer education, however, both models exceed the recommended sixth-grade reading level, indicating suboptimal optimization for general health literacy.

Burak Altunpak · 0 citations
Aug 2026

Performance of Large Language Models in Oral Cancer Patient Education: An Evaluation of Reliability, Readability, and Patient Communication Quality

Evaluating the reliability and readability of the responses generated by four mainstream LLMs to questions related to oral cancer found no model showed consistently high performance across all dimensions or met recommended readability standards.

Bo Zhang, Weidi Shi, Ying Zhang · 0 citations
Open access Aug 2026

Mapping Gaps and Improvement Targets in Large Language Model-Generated Melanoma Patient Education in a Non-English Setting

How well large language models (LLM) handle Turkish melanoma patient education varies widely from one model to the next, and findings suggest that LLM-generated Turkish melanoma materials may be useful as preliminary educational drafts.

Nıyazı Çetın, A. Atılan · 0 citations
Review Open access Aug 2026

A locally deployed large language model for pathology-informed and nurse-reviewed communication support in bladder cancer immunotherapy

Background Immunotherapy plays an important role in bladder cancer care, requiring ongoing patient education, symptom monitoring, and communication of pathology- and biomarker-related information. Locally deployed large language models (LLMs) may support these nurse-led activities, but their safety and clinical usabili...

Suqing Diao, Jun Zhao, Xu-Zhong Liu · 0 citations
Open access Sep 2026

Comparative evaluation of ChatGPT and gemini responses to patient-oriented questions on breast cancer

ChatGPT’s higher scores and shorter, more focused responses indicate that it may be a more efficient tool for addressing breast cancer–related patient questions, and Gemini’s acceptable overall performance demonstrates acceptable overall performance.

Yeliz Yılmaz Bozok, N. Acar, M. Atahan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.