Skip to content
#generative ai Open access

Generative AI responses to exercise-related questions asked by patients with multiple sclerosis: An evaluation of quality, accuracy, and readability.

Sep 2026 · Multiple Sclerosis and Related Disorders · Vol 115, pp. 107937 · 0 citations · 30 references
Medicine

TL;DR

ChatGPT-generated responses demonstrated higher quality, accuracy, and reliability than those generated by Gemini, Nonetheless, patient accessibility could be limited by the poor readability metrics observed in both tools.

Abstract

Objective

This study aims to compare the quality, accuracy, reliability, and readability of responses produced by ChatGPT-5.4 Thinking and Google Gemini 3 Flash regarding frequently asked questions by multiple sclerosis (MS) patients about exercise.

Method

A total of 75 questions were evaluated. Expert physiotherapists analysed the responses utilizing the Global Quality Score (GQS), Modified DISCERN (mDISCERN), Likert Accuracy Scale, and the Flesch Reading Ease (FRE).

Results

While 61.3% of ChatGPT responses were classified as high quality, this rate was 24% for Gemini (p < 0.001). ChatGPT-generated responses demonstrated significantly higher overall accuracy and mDISCERN scores than those generated by Gemini (p < 0.001). Categorical analyses revealed that the responses generated by ChatGPT received higher accuracy in the domains of exercise planning, safety and risk management, as well as participation and self-management. Additionally, ChatGPT-generated responses received significantly higher mDISCERN scores in the exercise planning and safety categories. No significant difference was observed between the two models regarding overall FRE scores (p > 0.05). The mean FRE scores for ChatGPT and Gemini were 41.70 and 38.21, respectively, with responses from both models classified at a "difficult" readability level.

Conclusion

Within the scope of this study, ChatGPT-generated responses demonstrated higher quality, accuracy, and reliability than those generated by Gemini. Nonetheless, patient accessibility could be limited by the poor readability metrics observed in both tools. While AI-based chatbots may show promise in reinforcing patient education for MS populations, they must not substitute specialized medical professionals during clinical decision-making.

Read PDF

Similar papers

Review Open access Aug 2026

Quality, readability, and patient safety of ChatGPT-generated responses to fall-related questions in older adults: a multidisciplinary evaluation

While ChatGPT provided generally acceptable clinical information, variability in readability and expert ratings raises patient safety concerns and AI-generated health content should be reviewed and tailored to older adults' health literacy needs before clinical use.

Merve Arı, N. Ilçin, Hatice Yağcıoğlu et al. · 0 citations
Open access Sep 2026

Artificial Intelligence in Bariatric Patient Education: A Multi-rater Evaluation of Reliability, Readability, and Clinical Validity of ChatGPT 5.2

ChatGPT 5.2 is a valuable AI–assisted chatbot that facilitates patient education by providing responses regarding sleeve gastrectomy that are generally accurate and acceptable, but the categorization of 4–16% of the responses as "Incorrect," the overall difficult readability levels, and the significant variability obse...

Furkan Türkoğlu, Elif Nur Gencer, Emre Erdoğan · 0 citations
Open access Sep 2026

Accuracy and limitations of artificial intelligence chatbots in answering patient questions on scoliosis surgery.

Objective The aim of this study was to analyze the accuracy, reliability, and quality of the content of responses provided by artificial intelligence (AI)-based chatbots to frequently asked questions related to scoliosis surgery. Methods A set of 25 questions related to diagnosis, treatment options, surgical risks, a...

G. Alibakan, Y. Sulek · 0 citations

Related blog posts

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.