BACKGROUND
Robotic-assisted total knee arthroplasty (rTKA) is increasingly used because of its surgical precision. However, inconsistent outcomes and high costs often lead patients to seek additional information from artificial intelligence (AI) tools. Large language models (LLMs) such as ChatGPT-4o, Gemini-2.5-Flash, and DeepSeek-V3 are commonly used, but their reliability and readability in orthopaedics remain unclear.
OBJECTIVES
To compare the reliability, usefulness, quality, and readability of responses to common patient questions about rTKA generated by leading LLMs.
METHODS
Three LLMs answered 20 frequently asked patient questions (n = 20) identified through Google Trends and expert validation. Three orthopaedic specialists (n = 3) evaluated reliability, usefulness, and overall quality using validated scales, while readability was assessed with standard indices.
RESULTS
Inter-rater reliability was good to excellent (ICC = 0.728-0.879). Gemini-2.5-Flash achieved significantly higher reliability and usefulness scores than ChatGPT-4o and DeepSeek-V3 (all p < 0.05). ChatGPT-4o and DeepSeek-V3 produced more readable but less accurate content, revealing an inverse relationship between reliability and readability.
CONCLUSIONS
Gemini-2.5-Flash provided the most reliable responses, highlighting the need for supervised integration of LLMs in patient education.
Mehmet Utku Çiftçi, A. Koluman, Ebru Aloğlu Çiftçi et al.· Knee (Oxford)· 0 citations
Patients with hallux valgus increasingly seek health information through consumer-facing artificial intelligence (AI)–driven patient education tools, particularly large language model–based conversational agents. Although these tools offer rapid and accessible responses, concerns remain regarding the reliability, usefulness, overall quality, and readability of AI-generated patient education materials. Evidence specifically evaluating AI-generated patient education for hallux valgus, a condition strongly influenced by patient expectations and treatment preferences, remains limited.
This cross-sectional comparative study evaluated the performance of three large language models—ChatGPT-4o, Gemini-2.5-Flash, and DeepSeek-V3—in responding to 20 patient-centered questions related to hallux valgus. Questions were developed using AI-assisted question generation and publicly available Google Trends search patterns and categorized into four clinical domains. AI-generated responses were anonymized and independently assessed by three orthopaedic surgeons for reliability, usefulness, and overall quality using 7-point Likert-based reliability and usefulness scales and the Global Quality Scale (GQS). Readability was analyzed using six standardized indices. Inter-rater agreement and between-model comparisons were statistically evaluated.
Gemini-2.5-Flash demonstrated modestly higher overall reliability, particularly in questions related to etiology and clinical presentation. DeepSeek-V3 achieved higher usefulness scores in the long-term outcomes and quality-of-life domain and produced significantly more readable content, as reflected by higher Flesch Reading Ease scores and lower grade-level indices. In contrast, Gemini-2.5-Flash generated linguistically more complex responses requiring higher educational levels for comprehension. Overall usefulness and global quality scores did not differ significantly among models. Qualitative review also identified occasional examples of oversimplified or potentially misleading information. Despite these differences, several readability metrics exceeded recommended patient health-literacy thresholds.
Contemporary AI-based conversational agents can provide patient-oriented information regarding hallux valgus with variable reliability and readability characteristics, although statistically significant differences were observed across models. A trade-off between factual accuracy and linguistic accessibility was observed. AI tools should therefore be regarded as adjuncts to, rather than replacements for, clinician-led patient education. Awareness of AI limitations and appropriate clinical guidance remain essential to ensure safe, accurate, and patient-centered information delivery.
Not applicable.
A. Koluman, Ebru Aloğlu Çiftçi, Mehmet Utku Çiftçi et al.· BMC Medical Informatics and...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.