Comparative quality, accuracy, and readability of large language model responses to patient questions about robotic-assisted total knee arthroplasty.
PURPOSE To compare the information quality, accuracy, and readability of patient-directed responses generated by large language models (LLMs), including ChatGPT-o3, ChatGPT-5.2, Gemini 3, and DeepSeek, regarding robotic-assisted total knee arthroplasty (RA-TKA). METHODS Thirty frequently asked patient questions were identified using LLM outputs and Google search queries. Responses were evaluated for information quality using the DISCERN and Quality Analysis of Medical Artificial Intelligence (QAMAI) instruments, for clinical accuracy using a 5-point ordinal rating scale, and for understandability and readability using the PEMAT Understandability and Flesch-Kincaid Reading Ease scores. RESULTS Median DISCERN scores were 46.0 (range, 35.0-50.0) for ChatGPT-o3, 45.75 (28.5-51.0) for ChatGPT-5.2, 43.75 (32.0-47.5) for Gemini 3, and 42.0 (32.0-50.0) for DeepSeek, with a significant overall difference among models (p < 0.001). The 5-point clinical accuracy scores were similar across models (median 4.0; p = 0.636). Median QAMAI scores were 23.0 for all four models, without a significant between-model difference (p = 0.462). PEMAT Understandability scores differed significantly among models (p < 0.001), with median scores of 90.0 for ChatGPT-o3 and ChatGPT-5.2, 88.0 for Gemini 3, and 85.0 for DeepSeek. Flesch-Kincaid Reading Ease scores also differed significantly (p < 0.001); Gemini 3 demonstrated higher readability than both ChatGPT models, whereas DeepSeek demonstrated higher readability than ChatGPT-o3. CONCLUSION The evaluated LLMs demonstrated generally acceptable clinical accuracy but differed across measures of written information quality, understandability, and readability. No significant difference was detected using QAMAI. Although Gemini 3 and DeepSeek demonstrated greater readability in selected comparisons, median responses across all models remained above recommended patient-education reading levels. LLM-generated responses should therefore be regarded as supplementary rather than standalone sources of patient information regarding RA-TKA.