Skip to content
Open access

MOST FREQUENTLY ASKED QUESTIONS BY OLDER ADULTS IN GERIATRIC REHABILITATION: EVALUATING LARGE LANGUAGE MODELS AS A SOURCE OF INFORMATION

Uğur Sözlü Selim Mahmut Günay Sevda Demir Türe Gül Özkazanç
2026 · Turkish Journal of Geriatrics · Vol 29 · 0 citations · 23 references

TL;DR

Although all three large language models generally produced reliable content, ChatGPT-5.2 and DeepSeek V3.2 showed stronger clinical accuracy performance, and because of the risk of incorrect information being generated, the use of large language models by the older population should preferably be done under expert supervision.

Abstract

Introduction: Older adults in geriatric rehabilitation are increasingly turning to internet-based resources and large language models for healthrelated information outside of clinical follow-up. The aim of this study was to compare the reliability, clinical accuracy, quality, usefulness, and readability of ChatGPT-5.2, Gemini 3, and DeepSeek V3.2 responses to patient questions regarding geriatric rehabilitation. Materials and Method: In this cross-sectional comparative content analysis, 24 predefined questions on geriatric rehabilitation were developed from YouTube comments, relevant literature, and clinical experience. Each question was submitted to three large language models under standardized conditions. Anonymized responses were independently evaluated by two experienced physiotherapists for reliability, clinical accuracy, quality, and usefulness, with disagreements resolved by consensus with a specialist physician. Readability was assessed using the Flesch Reading Ease scores. Results: No statistically significant difference was found among the three large language models in terms of reliability scores (p = 0.097). However, significant differences were observed for clinical accuracy, quality, usefulness, readability, and text characteristics (all p = 0.001). ChatGPT-5.2 and DeepSeek V3.2 showed the highest clinical accuracy scores, while ChatGPT-5.2 was superior in terms of quality and usefulness. For readability, ChatGPT-5.2 and Gemini 3 outperformed DeepSeek V3.2. Conclusion: Although all three large language models generally produced reliable content, ChatGPT-5.2 and DeepSeek V3.2 showed stronger clinical accuracy performance. Nevertheless, because of the risk of incorrect information being generated, the use of large language models by the older population should preferably be done under expert supervision. Keywords: Geriatrics; Rehabilitation; Artificial Intelligence; Patient Education as Topic; Health Literacy.

Read PDF

Similar papers

Review Open access Aug 2026

Quality, readability, and patient safety of ChatGPT-generated responses to fall-related questions in older adults: a multidisciplinary evaluation

SUMMARY OBJECTIVE: Older adults increasingly use artificial intelligence-based tools to obtain health information. Although artificial intelligence chatbots such as ChatGPT may enhance access, the quality, readability, and patient safety of fall-prevention information remain uncertain. This study aimed to evaluate the quality, readability, and patient safety implications of ChatGPT-generated responses to common questions about fall risk and home safety in older adults. METHODS: Ten frequently asked fall-related questions were submitted to ChatGPT (version 5.2). Responses were independently assessed by a multidisciplinary panel including physiotherapists, a geriatrician, a physical medicine and rehabilitation physician, an occupational therapist, and an orthopedic specialist. Quality was evaluated using the Mika classification. Readability was measured with the Flesch-Kincaid Grade Level. Interrater reliability was analyzed using a two-way random-effects intraclass correlation coefficient model with absolute agreement (intraclass correlation coefficient [2,k]). RESULTS: Three responses were rated as "excellent," while seven responses were rated as "satisfactory requiring minimal clarification." No response received a rating corresponding to "moderately satisfactory" or "unsatisfactory." The mean Flesch-Kincaid Grade Level was 8.4 (range 4.3–11.9). Five responses exceeded the readability levels commonly recommended for patient education materials. Interrater reliability demonstrated fair agreement (intraclass correlation coefficient [2,k]=0.72; 95%CI 0.64–0.80). CONCLUSION: While ChatGPT provided generally acceptable clinical information, variability in readability and expert ratings raises patient safety concerns. AI-generated health content should be reviewed and tailored to older adults’ health literacy needs before clinical use.

Merve Arı, N. Ilçin, Hatice Yağcıoğlu et al. · 0 citations
#small language model Preprint Aug 2026

Performance of a domain-specific large language model in answering patient questions in psychiatry

MIND was able to answer many questions about escitalopram in a manner deemed accurate, complete, and safe by psychiatrists the majority of the time, and represents a step towards building safe LLM systems to enhance patient education in psychiatry.

Alexander J. Hish, A. Nagendran, S. Compton · 0 citations
Review Open access Sep 2026

Comparing Chinese-language large language models for caregiver questions about developmental dysplasia of the hip

Large language models (LLMs) are increasingly used for caregiver-facing health information, but their reliability in Chinese-language pediatric orthopaedics remains uncertain. This study evaluated whether responses to developmental dysplasia of the hip (DDH) questions were clinically accurate, aligned with Chinese guidance, and educationally usable. We compared ChatGPT-4o and DeepSeek-R1 using 53 Chinese-language DDH questions, including 31 caregiver-oriented frequently asked questions and 22 guideline-derived items. Each model generated one response per question using a standardized single-turn, five-sentence prompt. Six blinded pediatric orthopaedic surgeons rated clinical accuracy and guideline concordance. Paired model comparisons, inter-rater reliability, and exploratory formula-based readability were assessed. DeepSeek-R1 had higher clinical accuracy than ChatGPT-4o across 31 caregiver-oriented questions (mean question-level score 4.70 [SD 0.15] vs. 3.69 [0.28]; P  < 0.001) and higher guideline concordance across 22 items (4.44 [0.22] vs. 3.82 [0.33]; P  < 0.001). Average-rating reliability was good for clinical accuracy (ICC(2,6) = 0.826, 95% bootstrap CI 0.770−0.863) and moderate for guideline concordance (ICC(2,6) = 0.519, 95% CI 0.377−0.621). Formula-based readability findings were mixed: unadjusted tests indicated lower FKGL, FOG, and SMOG values for ChatGPT-4o, but only FOG and SMOG remained statistically significant after Holm adjustment. Surgery-related observations were exploratory and descriptive. Under the tested single-query and five-sentence conditions, DeepSeek-R1 achieved higher expert-rated clinical accuracy and Chinese-guideline concordance. The findings describe specific model-platform configurations rather than a permanent model ranking. Professionally reviewed LLM responses may help clinicians reinforce routine DDH education, but individualized diagnostic and treatment guidance, especially for surgery-related questions, should remain clinician-led. Caregivers should use chatbot information only as a supplementary resource.

Unknown authors · 0 citations
Open access Aug 2026

Evaluating large language models in patient education: a comparative analysis addressing frequently asked questions in peri-acetabular osteotomy.

INTRODUCTION Large language models (LLMs) are increasingly used as sources of information across many fields, including healthcare. As patients turn to these models for health-related queries, evaluating the accuracy and reliability of their responses is essential. Peri-acetabular osteotomy (PAO) is performed on younger patients - a group more likely to use digital tools like LLMs for health information. This study assesses the accuracy and readability performance of two leading LLMs, ChatGPT and Google Gemini, in addressing common patient questions on PAO. METHODS A panel of fellowship-trained PAO surgeons created ten commonly asked patient questions based on real-world experience. Responses from each LLM were assessed by the same three surgeons, blinded to response origin, using a 5-point Likert scale to evaluate clarity, accuracy, and completeness. Readability was measured with Flesch-Kincaid Reading scores. RESULTS ChatGPT outperformed Gemini with an average score of 4.17 vs 3.13 (t = -3.08, p = 0.006). ChatGPT's responses were often rated higher for completeness and clarity, particularly in areas needing detailed explanation, and usually required minimal clarification. Gemini sometimes lacked specificity or included minor inaccuracies that reduced its perceived reliability. Both LLMs produced responses with similar "difficult" Flesch Reading Ease scores. CONCLUSIONS There may be significant differences in how effectively LLMs support patients with surgical queries. ChatGPT more consistently met expert standards for clarity and thoroughness. As LLM usage expands, ChatGPT may aid patient education on hip surgery, supporting consultations, informed decisions and postoperative guidance.

TP Davis, B. Guevel, K. Logishetty et al. · 0 citations
Open access Aug 2026

Are large language models such as ChatGPT, capable of supporting patients and general practitioners after spine surgery?

LLMs can support communication and education following spine surgery when used with structured prompting when used with structured prompting and ChatGPT and Claude showed the highest correctness and completeness, particularly for practitioner-directed answers.

S. Wegmann, T. Rosenkranz, Philipp Egenolf et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.