Although all three large language models generally produced reliable content, ChatGPT-5.2 and DeepSeek V3.2 showed stronger clinical accuracy performance, and because of the risk of incorrect information being generated, the use of large language models by the older population should preferably be done under expert supervision.
Abstract
Introduction: Older adults in geriatric rehabilitation are increasingly
turning to internet-based resources and large language models for healthrelated information outside of clinical follow-up. The aim of this study was to
compare the reliability, clinical accuracy, quality, usefulness, and readability of
ChatGPT-5.2, Gemini 3, and DeepSeek V3.2 responses to patient questions
regarding geriatric rehabilitation.
Materials and Method: In this cross-sectional comparative content analysis,
24 predefined questions on geriatric rehabilitation were developed from
YouTube comments, relevant literature, and clinical experience. Each question
was submitted to three large language models under standardized conditions.
Anonymized responses were independently evaluated by two experienced
physiotherapists for reliability, clinical accuracy, quality, and usefulness, with
disagreements resolved by consensus with a specialist physician. Readability
was assessed using the Flesch Reading Ease scores.
Results: No statistically significant difference was found among the three
large language models in terms of reliability scores (p = 0.097). However,
significant differences were observed for clinical accuracy, quality, usefulness,
readability, and text characteristics (all p = 0.001). ChatGPT-5.2 and DeepSeek
V3.2 showed the highest clinical accuracy scores, while ChatGPT-5.2 was
superior in terms of quality and usefulness. For readability, ChatGPT-5.2 and
Gemini 3 outperformed DeepSeek V3.2.
Conclusion: Although all three large language models generally produced
reliable content, ChatGPT-5.2 and DeepSeek V3.2 showed stronger clinical
accuracy performance. Nevertheless, because of the risk of incorrect information
being generated, the use of large language models by the older population
should preferably be done under expert supervision.
Keywords: Geriatrics; Rehabilitation; Artificial Intelligence; Patient
Education as Topic; Health Literacy.
SUMMARY OBJECTIVE: Older adults increasingly use artificial intelligence-based tools to obtain health information. Although artificial intelligence chatbots such as ChatGPT may enhance access, the quality, readability, and patient safety of fall-prevention information remain uncertain. This study aimed to evaluate the quality, readability, and patient safety implications of ChatGPT-generated responses to common questions about fall risk and home safety in older adults. METHODS: Ten frequently asked fall-related questions were submitted to ChatGPT (version 5.2). Responses were independently assessed by a multidisciplinary panel including physiotherapists, a geriatrician, a physical medicine and rehabilitation physician, an occupational therapist, and an orthopedic specialist. Quality was evaluated using the Mika classification. Readability was measured with the Flesch-Kincaid Grade Level. Interrater reliability was analyzed using a two-way random-effects intraclass correlation coefficient model with absolute agreement (intraclass correlation coefficient [2,k]). RESULTS: Three responses were rated as "excellent," while seven responses were rated as "satisfactory requiring minimal clarification." No response received a rating corresponding to "moderately satisfactory" or "unsatisfactory." The mean Flesch-Kincaid Grade Level was 8.4 (range 4.3–11.9). Five responses exceeded the readability levels commonly recommended for patient education materials. Interrater reliability demonstrated fair agreement (intraclass correlation coefficient [2,k]=0.72; 95%CI 0.64–0.80). CONCLUSION: While ChatGPT provided generally acceptable clinical information, variability in readability and expert ratings raises patient safety concerns. AI-generated health content should be reviewed and tailored to older adults’ health literacy needs before clinical use.
Merve Arı, N. Ilçin, Hatice Yağcıoğlu et al.· Revista da Associação Médica...· 0 citations
MIND was able to answer many questions about escitalopram in a manner deemed accurate, complete, and safe by psychiatrists the majority of the time, and represents a step towards building safe LLM systems to enhance patient education in psychiatry.
Alexander J. Hish, A. Nagendran, S. Compton· 0 citations
Large language models (LLMs) are increasingly used for caregiver-facing health information, but their reliability in Chinese-language pediatric orthopaedics remains uncertain. This study evaluated whether responses to developmental dysplasia of the hip (DDH) questions were clinically accurate, aligned with Chinese guidance, and educationally usable.
We compared ChatGPT-4o and DeepSeek-R1 using 53 Chinese-language DDH questions, including 31 caregiver-oriented frequently asked questions and 22 guideline-derived items. Each model generated one response per question using a standardized single-turn, five-sentence prompt. Six blinded pediatric orthopaedic surgeons rated clinical accuracy and guideline concordance. Paired model comparisons, inter-rater reliability, and exploratory formula-based readability were assessed.
DeepSeek-R1 had higher clinical accuracy than ChatGPT-4o across 31 caregiver-oriented questions (mean question-level score 4.70 [SD 0.15] vs. 3.69 [0.28];
P
< 0.001) and higher guideline concordance across 22 items (4.44 [0.22] vs. 3.82 [0.33];
P
< 0.001). Average-rating reliability was good for clinical accuracy (ICC(2,6) = 0.826, 95% bootstrap CI 0.770−0.863) and moderate for guideline concordance (ICC(2,6) = 0.519, 95% CI 0.377−0.621). Formula-based readability findings were mixed: unadjusted tests indicated lower FKGL, FOG, and SMOG values for ChatGPT-4o, but only FOG and SMOG remained statistically significant after Holm adjustment. Surgery-related observations were exploratory and descriptive.
Under the tested single-query and five-sentence conditions, DeepSeek-R1 achieved higher expert-rated clinical accuracy and Chinese-guideline concordance. The findings describe specific model-platform configurations rather than a permanent model ranking. Professionally reviewed LLM responses may help clinicians reinforce routine DDH education, but individualized diagnostic and treatment guidance, especially for surgery-related questions, should remain clinician-led. Caregivers should use chatbot information only as a supplementary resource.
Unknown authors· Frontiers in Pediatrics· 0 citations
INTRODUCTION
Large language models (LLMs) are increasingly used as sources of information across many fields, including healthcare. As patients turn to these models for health-related queries, evaluating the accuracy and reliability of their responses is essential. Peri-acetabular osteotomy (PAO) is performed on younger patients - a group more likely to use digital tools like LLMs for health information. This study assesses the accuracy and readability performance of two leading LLMs, ChatGPT and Google Gemini, in addressing common patient questions on PAO.
METHODS
A panel of fellowship-trained PAO surgeons created ten commonly asked patient questions based on real-world experience. Responses from each LLM were assessed by the same three surgeons, blinded to response origin, using a 5-point Likert scale to evaluate clarity, accuracy, and completeness. Readability was measured with Flesch-Kincaid Reading scores.
RESULTS
ChatGPT outperformed Gemini with an average score of 4.17 vs 3.13 (t = -3.08, p = 0.006). ChatGPT's responses were often rated higher for completeness and clarity, particularly in areas needing detailed explanation, and usually required minimal clarification. Gemini sometimes lacked specificity or included minor inaccuracies that reduced its perceived reliability. Both LLMs produced responses with similar "difficult" Flesch Reading Ease scores.
CONCLUSIONS
There may be significant differences in how effectively LLMs support patients with surgical queries. ChatGPT more consistently met expert standards for clarity and thoroughness. As LLM usage expands, ChatGPT may aid patient education on hip surgery, supporting consultations, informed decisions and postoperative guidance.
TP Davis, B. Guevel, K. Logishetty et al.· Annals of the Royal College...· 0 citations
LLMs can support communication and education following spine surgery when used with structured prompting when used with structured prompting and ChatGPT and Claude showed the highest correctness and completeness, particularly for practitioner-directed answers.
S. Wegmann, T. Rosenkranz, Philipp Egenolf et al.· European spine journal· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.