Skip to content
Open access

Comparative evaluation of ChatGPT and gemini responses to patient-oriented questions on breast cancer

Sep 2026 · PLoS ONE · Vol 21, pp. e0358325 · 0 citations · 26 references
Medicine

TL;DR

ChatGPT’s higher scores and shorter, more focused responses indicate that it may be a more efficient tool for addressing breast cancer–related patient questions, and Gemini’s acceptable overall performance demonstrates acceptable overall performance.

Abstract

Background Breast cancer, the most common malignancy among women, remains a major global public health concern. With the rapid growth of artificial intelligence–based language models, it is essential to evaluate their potential roles in patient education. This study compared ChatGPT and Gemini in responding to patient-oriented questions about breast cancer and its surgical treatment regarding scientific accuracy, clarity, and unnecessary detail. Methods Forty frequently asked questions were collected from Turkish online sources. Both models were queried under identical conditions on August 13, 2025, and the responses were anonymized for blinded evaluation. Four general surgeons experienced in breast surgery independently assessed each response using a five-point Likert scale across three domains: scientific accuracy, clarity, and unnecessary detail. Results Response lengths were recorded and compared. ChatGPT achieved significantly higher median scores than Gemini in scientific accuracy [4.75 (3.50–5.00) vs. 4.25 (3.50–5.00); p < 0.001], clarity [4.75 (3.50–5.00) vs. 4.25 (3.25–5.00); p = 0.005], and unnecessary detail [5.00 (5.00–5.00) vs. 4.50 (3.50–5.00); p < 0.001]. The overall median score was 4.83 (4.08–5.00) for ChatGPT and 4.33 (3.50–4.92) for Gemini (p < 0.001). Gemini’s responses were significantly longer (244 ± 84 vs. 170 ± 47 words; p < 0.001). Conclusion Both models demonstrated acceptable overall performance. However, ChatGPT’s higher scores and shorter, more focused responses indicate that it may be a more efficient tool for addressing breast cancer–related patient questions. In their current form, these models should not be used for diagnostic or clinical decision-making purposes. AI-generated health information should be used only under expert supervision and for educational or supportive purposes.

Read PDF

Similar papers

Sep 2026

ChatGPT 4.o Responses to Commonly Asked Colon and Rectal Cancer Questions From a Patient Perspective.

IntroductionPatients are increasingly turning to large language models (LLMs) such as ChatGPT for medical guidance, prompting concerns around medical competence. This study assesses the proficiency of ChatGPT 4.o in answering patient questions regarding the surgical management of colon and rectal cancer.MethodsChatGPT...

Alyssa Habermann, Tania Gupta, S. Cohen et al. · 0 citations
Open access Aug 2026

Evaluation of AI Chatbot Responses to Pediatric Urology Frequently Asked Questions.

OBJECTIVE To evaluate the quality of responses from four publicly available LLMs (ChatGPT-4o, Claude 3.7, Gemini 2.5, and Copilot) to frequently asked questions (FAQs) in pediatric urology. METHODS FAQs were generated using standardized prompts and submitted to each LLM using parent-centered instructions. Two board-c...

Najva Mazhari, Andrew Freedman, Nadine A. Friedrich et al. · 0 citations
Open access Aug 2026

Safety and quality of public chatbots for lung cancer prognostic information: a comparative evaluation

Public-facing chatbots may support general patient education but should not replace individualized clinician-led prognostic communication as public-facing chatbots differed substantially in safety, reliability, communication quality, and readability.

Yan-Ru Jiang, Qian-Yun Wang, Liang Zheng et al. · 0 citations
Aug 2026

Performance of Large Language Models in Oral Cancer Patient Education: An Evaluation of Reliability, Readability, and Patient Communication Quality

Evaluating the reliability and readability of the responses generated by four mainstream LLMs to questions related to oral cancer found no model showed consistently high performance across all dimensions or met recommended readability standards.

Bo Zhang, Weidi Shi, Ying Zhang · 0 citations
Sep 2026

Large Language Models for Breast Cancer Education: A Comparative Analysis of Quality, Reliability and Readability.

Gemini significantly outperforms ChatGPT in response quality, reliability, and linguistic accessibility for breast cancer education, however, both models exceed the recommended sixth-grade reading level, indicating suboptimal optimization for general health literacy.

Burak Altunpak · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.