Skip to content

Similar papers

Aug 2026

Comparative performance of artificial intelligence chatbots in patient education for robot-assisted radical prostatectomy: quality, transparency and readability

Although some models provided comparatively stronger patient education, none consistently produced sufficiently accessible information on robot-assisted radical prostatectomy, and these tools may supplement, but should not replace, individualized counseling by urologists and multidisciplinary prostate cancer teams.

Yang Liu, Rong-Kang Li, Peng Yu et al. · 0 citations
Open access Feb 2026

ChatGPT as a source of surgical information: Evaluation of responses to patient questions on hallux rigidus fusion

Background Patients are increasingly turning to online resources and artificial intelligence (AI)-based tools to obtain information about orthopedic conditions and surgical options. Large language models, such as ChatGPT, are becoming prominent in patient education; however, their reliability and readability remain uncertain. This study evaluated the quality and readability of responses generated by ChatGPT-4o and ChatGPT-5 to frequently asked patient questions regarding hallux rigidus fusion surgery. Methods Twenty commonly asked patient questions were compiled and presented to ChatGPT-4o and ChatGPT-5. Readability was assessed using the Flesch–Kincaid Grade Level, Gunning Fog, Coleman–Liau, and Simple Measure of Gobbledygook indices. Quality was evaluated with the DISCERN tool, response accuracy scores, and Journal of the American Medical Association (JAMA) criteria. Interrater agreement was measured using the Intraclass Correlation Coefficient (ICC). Results ChatGPT-4o generated longer responses (802 vs. 242 words; p<0.001) with slightly higher readability grade levels (10.81 vs. 10.37; p=0.031). Accuracy (2.00 vs. 1.85; p=0.323) and DISCERN scores (49.35 vs. 48.93; p=0.747) showed no significant differences. All responses received a JAMA score of 0 due to the absence of citations, authorship, or transparency indicators. Interrater reliability indicated moderate to good agreement (ICC: 0.68–0.80). Conclusion ChatGPT-4o and ChatGPT-5 provide generally satisfactory yet non-comprehensive, limited-quality information at a level above tenth-grade regarding hallux rigidus fusion surgery. Although linguistically coherent, responses lack evidence-based detail and individualized guidance. These models may supplement, but cannot replace, expert orthopedic counseling. Ensuring physician oversight and integrating validated, updated clinical content remain essential for safe implementation of AI-generated patient information.

Kamil Balaban, Mehmet Batu Ertan, Mahmut Kalem · 0 citations
Open access Aug 2026

Quality and readability of AI Chatbot responses to frequently asked questions from patients undergoing progressive collapsing foot deformity surgery: a comparative study of ChatGPT, Perplexity, and Gemini.

BACKGROUND Patients increasingly turn to AI chatbots for medical information, including before complex orthopaedic procedures such as progressive collapsing foot deformity (PCFD) surgery. Whether these tools deliver content of sufficient quality and accessibility for preoperative patient education remains unclear, particularly across competing platforms. This study addressed three questions: (1) Do ChatGPT, Perplexity AI and Google Gemini differ in the accuracy, comprehensiveness and clarity of their responses to PCFD-related patient questions? (2) Do these platforms produce content meeting recommended readability thresholds for patient education? (3) Does the level of agreement among blinded foot and ankle surgeons rating the quality of chatbot responses vary depending on the platform used? HYPOTHESIS The three AI chatbot platforms produce responses of comparable accuracy but differ significantly in readability, with none reaching the recommended readability thresholds for patient education materials. PATIENTS AND METHODS Cross-sectional comparative study. Twenty frequently asked questions regarding PCFD, covering disease understanding, conservative management, surgical planning and postoperative recovery, were submitted verbatim to ChatGPT (GPT-4o mini), Perplexity AI and Google Gemini (free versions, March 25, 2026). The 60 resulting responses were rated by three blinded foot and ankle surgeons on three 5-point Likert scales (accuracy, comprehensiveness, clarity). Readability was assessed using the Flesch Reading Ease (FRE) and Flesch-Kincaid Grade Level (FKGL). Inter-rater agreement used Kendall's W; differences between platforms were analysed using Kruskal-Wallis tests with Bonferroni-corrected pairwise comparisons. RESULTS All platforms produced responses rated accurate to very accurate. Perplexity achieved significantly higher accuracy than ChatGPT (p = 0.0003) and higher accuracy and clarity than Gemini (p = 0.0065 and p = 0.0075). No platform reached the recommended FRE ≥ 60 or FKGL ≤ 6 thresholds: median FKGL ranged from 13.1 (ChatGPT) to 21.4 (Perplexity), with ChatGPT producing the most readable and Perplexity the least readable content (p < 0.001). Inter-rater agreement was fair to substantial across platforms, lowest for Perplexity. DISCUSSION AI chatbots produce generally accurate baseline information on PCFD surgery, with Perplexity showing significantly higher expert-rated accuracy and clarity than the other platforms-contrary to our hypothesis of comparable accuracy-while readability remains uniformly inadequate for all platforms, as hypothesized. These tools may serve as a supplementary source of information, but their inadequate readability suggests they are not yet suited to replace tailored, surgeon-led patient education. LEVEL OF EVIDENCE III; cross-sectional comparative study.

L. Micicoi, J. Brué, Saumith Menon et al. · 0 citations
Aug 2026

Quality of AI-Generated Patient Education for Pre- and Post-Operative Tracheostomy Care.

OBJECTIVE To evaluate the accuracy, completeness, clarity, source transparency, and readability of leading AI chatbot responses to patient questions about tracheostomy and to determine whether AI tools can reliably support patient education where high-quality guidance is critical for safety. STUDY DESIGN Cross-sectional content analysis. SETTING Virtual study environment using publicly accessible AI platforms, with expert evaluation conducted via Qualtrics-based distribution. METHODS Twelve frequently asked questions about tracheostomy care were identified using search-listening tools and clinician input, then submitted to 5 AI chatbots - ChatGPT4, Google Gemini 2.0, Microsoft Copilot, DeepSeek V3, and Grok 3 - and to a senior laryngologist. Three blinded laryngologists independently evaluated each response using the Quality Analysis of Medical Artificial Intelligence instrument. Readability was assessed using nine metrics. RESULTS Gemini 2.0 achieved significantly higher completeness scores than physician responses (P < .001), with DeepSeek and Grok 3 (P < .05) also outperforming (P < .05). Accuracy did not differ significantly between AI- and expert-generated responses. On average, the AI models outperformed physician in clarity, completeness, and usefulness based on QAMAI scoring (P < .05). All AI and expert responses exceeded the NIH-recommended 6th-grade reading level, ranging from 10th-13th grade (P < .001). Inter-rater reliability was 78%. CONCLUSION AI chatbots can generate accurate and comprehensive responses to common tracheostomy care questions, demonstrating potential to support patient education. However, they continue to lack guaranteed, verifiable sourcing, and this study did not assess actual patient comprehension of the AI-generated responses. Future efforts should focus on adapting AI-generated education materials to meet health literacy standards and evaluating their direct impact on patient understanding and outcomes.

Keer Zhang, Lauran K. Evans, Desiree Delavary et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.