Skip to content

Large Language Models for Breast Cancer Education: A Comparative Analysis of Quality, Reliability and Readability.

Sep 2026 · Surgical Innovation · pp. 15533506261488588 · 0 citations · 40 references
Medicine

TL;DR

Gemini significantly outperforms ChatGPT in response quality, reliability, and linguistic accessibility for breast cancer education, however, both models exceed the recommended sixth-grade reading level, indicating suboptimal optimization for general health literacy.

Abstract

BackgroundPatients increasingly consult artificial intelligence (AI) tools for breast cancer information. While Large Language Models (LLMs) enhance information accessibility, their accuracy, reliability, and alignment with patient health literacy remain critical concerns. This study compared the quality, reliability, and readability of breast cancer-related responses generated by ChatGPT and Gemini.MethodsIn this cross-sectional study, conducted between March 20 and March 31, 2026, 40 questions spanning diagnosis, treatment, genetics, and follow-up were submitted to ChatGPT-5.3 and Gemini 3.0 Flash. Three independent surgeons evaluated the responses in a double-blinded manner using the modified DISCERN (mDISCERN) for reliability and Global Quality Score (GQS) for content quality assessment. Readability was assessed via Flesch Reading Ease (FRES), Flesch-Kincaid Grade Level (FKGL), Gunning Fog Index (GFI), and Simple Measure of Gobbledygook (SMOG) indices.ResultsGemini demonstrated statistically significant superiority over ChatGPT in both GQS (4.35 ± 0.30 vs 3.90 ± 0.29; P < .001) and mDISCERN (3.63 ± 0.58 vs 3.02 ± 0.48; P < .001) scores. In readability analysis, Gemini exhibited higher FRES (53.90 vs 45.13) and lower FKGL (9.31 vs 10.91) values, indicating enhanced patient accessibility (P < .05). For both models, the "Diagnosis" category yielded the highest readability, whereas "Treatment" scored the lowest. Inter-rater reliability for mDISCERN was moderate (ICC = 0.595).ConclusionsGemini significantly outperforms ChatGPT in response quality, reliability, and linguistic accessibility for breast cancer education. However, both models exceed the recommended sixth-grade reading level, indicating suboptimal optimization for general health literacy. While LLMs serve as promising auxiliary tools, expert supervision and cross-validation remain mandatory to ensure patient safety.

View source

Similar papers

Sep 2026

When Treatment Gets More Complicated: AI Readability Declines in Breast Cancer Patient Education.

BackgroundThe proliferation of Large Language Models necessitates evaluating their performance in communication accessibility. This study compares the linguistic architecture, readability, and automated psycholinguistic text properties of AI-generated breast cancer materials across varying clinical complexities.Methods...

Burak Altunpak · 0 citations
Aug 2026

Performance of Large Language Models in Oral Cancer Patient Education: An Evaluation of Reliability, Readability, and Patient Communication Quality

Evaluating the reliability and readability of the responses generated by four mainstream LLMs to questions related to oral cancer found no model showed consistently high performance across all dimensions or met recommended readability standards.

Bo Zhang, Weidi Shi, Ying Zhang · 0 citations
Open access Sep 2026

Performance evaluation of large language models in bladder cancer patient education Q&A: a cross-sectional study

Current mainstream LLMs demonstrate initial potential for generating educational content on bladder cancer, albeit with considerable heterogeneity across models, and advocate for a prudent, assistive role of LLMs in health communication under a human-AI collaborative model.

Dian Wan, You-Wen Li, Zheng Dong et al. · 0 citations
Review Open access Aug 2026

Evaluating AI-generated patient education materials for endometrial cancer surgery: a comparative analysis of response quality, reliability, and readability between ChatGPT and DeepSeek models

DeepSeek demonstrates a significant advantage in information reliability, particularly excelling in postoperative and follow-up management content, and ChatGPT shows a slight edge in the readability of surgical planning sections.

Ling Tian, Long-Yue Tang, Ming-Tao Yang et al. · 0 citations
Open access Aug 2026

Mapping Gaps and Improvement Targets in Large Language Model-Generated Melanoma Patient Education in a Non-English Setting

How well large language models (LLM) handle Turkish melanoma patient education varies widely from one model to the next, and findings suggest that LLM-generated Turkish melanoma materials may be useful as preliminary educational drafts.

Nıyazı Çetın, A. Atılan · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.