Skip to content

Author

Ebru Aloğlu Çiftçi

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

Large language models as sources of patient information on robotic knee arthroplasty: a comparative evaluation.

BACKGROUND Robotic-assisted total knee arthroplasty (rTKA) is increasingly used because of its surgical precision. However, inconsistent outcomes and high costs often lead patients to seek additional information from artificial intelligence (AI) tools. Large language models (LLMs) such as ChatGPT-4o, Gemini-2.5-Flash, and DeepSeek-V3 are commonly used, but their reliability and readability in orthopaedics remain unclear. OBJECTIVES To compare the reliability, usefulness, quality, and readability of responses to common patient questions about rTKA generated by leading LLMs. METHODS Three LLMs answered 20 frequently asked patient questions (n = 20) identified through Google Trends and expert validation. Three orthopaedic specialists (n = 3) evaluated reliability, usefulness, and overall quality using validated scales, while readability was assessed with standard indices. RESULTS Inter-rater reliability was good to excellent (ICC = 0.728-0.879). Gemini-2.5-Flash achieved significantly higher reliability and usefulness scores than ChatGPT-4o and DeepSeek-V3 (all p < 0.05). ChatGPT-4o and DeepSeek-V3 produced more readable but less accurate content, revealing an inverse relationship between reliability and readability. CONCLUSIONS Gemini-2.5-Flash provided the most reliable responses, highlighting the need for supervised integration of LLMs in patient education.

Mehmet Utku Çiftçi, A. Koluman, Ebru Aloğlu Çiftçi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.