Skip to content

Comparative performance of artificial intelligence chatbots in patient education for robot-assisted radical prostatectomy: quality, transparency and readability

Aug 2026 · Journal of Robotic Surgery · Vol 20 · 0 citations · 29 references
Medicine

TL;DR

Although some models provided comparatively stronger patient education, none consistently produced sufficiently accessible information on robot-assisted radical prostatectomy, and these tools may supplement, but should not replace, individualized counseling by urologists and multidisciplinary prostate cancer teams.

View source

Similar papers

Open access Sep 2026

Comparative quality, accuracy, and readability of large language model responses to patient questions about robotic-assisted total knee arthroplasty.

PURPOSE To compare the information quality, accuracy, and readability of patient-directed responses generated by large language models (LLMs), including ChatGPT-o3, ChatGPT-5.2, Gemini 3, and DeepSeek, regarding robotic-assisted total knee arthroplasty (RA-TKA). METHODS Thirty frequently asked patient questions were identified using LLM outputs and Google search queries. Responses were evaluated for information quality using the DISCERN and Quality Analysis of Medical Artificial Intelligence (QAMAI) instruments, for clinical accuracy using a 5-point ordinal rating scale, and for understandability and readability using the PEMAT Understandability and Flesch-Kincaid Reading Ease scores. RESULTS Median DISCERN scores were 46.0 (range, 35.0-50.0) for ChatGPT-o3, 45.75 (28.5-51.0) for ChatGPT-5.2, 43.75 (32.0-47.5) for Gemini 3, and 42.0 (32.0-50.0) for DeepSeek, with a significant overall difference among models (p < 0.001). The 5-point clinical accuracy scores were similar across models (median 4.0; p = 0.636). Median QAMAI scores were 23.0 for all four models, without a significant between-model difference (p = 0.462). PEMAT Understandability scores differed significantly among models (p < 0.001), with median scores of 90.0 for ChatGPT-o3 and ChatGPT-5.2, 88.0 for Gemini 3, and 85.0 for DeepSeek. Flesch-Kincaid Reading Ease scores also differed significantly (p < 0.001); Gemini 3 demonstrated higher readability than both ChatGPT models, whereas DeepSeek demonstrated higher readability than ChatGPT-o3. CONCLUSION The evaluated LLMs demonstrated generally acceptable clinical accuracy but differed across measures of written information quality, understandability, and readability. No significant difference was detected using QAMAI. Although Gemini 3 and DeepSeek demonstrated greater readability in selected comparisons, median responses across all models remained above recommended patient-education reading levels. LLM-generated responses should therefore be regarded as supplementary rather than standalone sources of patient information regarding RA-TKA.

U. Kolaç, Mazlum Veysel Sili, Orhan Mete Karademir et al. · 0 citations
Sep 2026

Guideline-Based Evaluation of Five Generative Artificial Intelligence Chatbot Platforms for Clinician-Oriented Questions on Open Temporomandibular Joint Surgery.

OBJECTIVE To compare the quality and readability of responses from five generative artificial intelligence chatbot platforms to clinician-oriented questions on open temporomandibular joint (TMJ) surgery against guideline-based reference answers. MATERIAL AND METHODS Forty questions across eight domains were submitted on 16 July 2026 to ChatGPT (GPT-5.5), Claude (Opus 4.8), Gemini (3.1 Pro), Grok (4) and Perplexity (Pro), each via its paid tier at default settings (200 responses). Two blinded oral and maxillofacial surgeons applied the Quality Analysis of Medical Artificial Intelligence (QAMAI) tool, the Global Quality Score (GQS) and a five-point overall quality rating using a priori key elements and written anchors. Readability was assessed with the Flesch-Kincaid Grade Level (FKGL) and Flesch Reading Ease Score (FRES); platforms were compared with the Friedman test and Bonferroni-corrected Wilcoxon post-hoc tests; inter-rater reliability used intraclass correlation coefficients (ICC). RESULTS Inter-rater reliability was good to excellent (average-measure ICC 0.93-0.99). All outcomes except clarity differed among platforms (P < .001). Perplexity achieved the highest QAMAI total (27.5 ± 1.4; Kendall's W = 0.75), largely through retrieval-based source provision; excluding this domain, Claude, Perplexity and Gemini converged. Claude had the highest GQS (4.5 ± 0.6), Gemini the highest overall rating (4.3 ± 0.7); Grok scored lowest. Claude and Gemini were least readable (median FKGL 22.7 and 26.1; FRES -11.6 and -7.5). Length did not explain scores within platforms. CONCLUSIONS Performance varied substantially across platforms and dimensions; no platform optimized all outcomes. These findings describe informational quality, not clinical safety or decision-making, and support specialist verification before clinical use.

Selin Gaş, Erdinç Sulukan, Büşra Korkmaz · 0 citations
Open access Jul 2026

Evaluation of artificial intelligence chatbots in informing patients about CBCT: A comparative study

Objective: This study aimed to explore common patient inquiries about Cone-Beam Computed Tomography (CBCT) and systematically assess the accuracy, quality, usefulness, and readability of outputs from four AI-based conversational platforms (ChatGPT-3.5, DeepSeek-V3.1, Gemini 1.5 Flash, and Microsoft Copilot Free).Materials and Methods: Common questions about CBCT were gathered from online sources and expert contributions and submitted to ChatGPT-3.5, DeepSeek, Gemini, and Copilot under standardized conditions. Outputs were independently evaluated by three specialists using CLEAR, mGQS, accuracy, usefulness, DISCERN, and readability metrics (FRE and FKGL).Results: Significant differences were observed among AI platforms regarding CLEAR scores (p < 0.05), with Gemini showing higher values than ChatGPT (p = 0.016). Across all four platforms, strong to a very strong positive correlations were found among CLEAR, mGQS, and Accuracy scores (r ≥ 0.77, p < 0.001), while these measures were strongly to very strongly and inversely correlated with Usefulness (r ≤ −0.77, p < 0.001). Flesch Reading Ease and Flesch-Kincaid Grade Level demonstrated a very strong negative correlations across platforms (r ranging from −0.85 to −0.93, p < 0.05). Within the Gemini group, Flesch-Kincaid Grade Level showed a moderate positive association with DISCERN scores (r = 0.496, p = 0.043).Conclusions: Gemini and DeepSeek produced more accurate and clearer responses. The readability of all AI systems was below the recommended level for patient education. These systems may support patient education but require expert validation.

Savaş Özarslantürk, Seval Ceylan Şen, Özlem Saraç Atagün et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.