Skip to content
Open access

Artificial Intelligence in Bariatric Patient Education: A Multi-rater Evaluation of Reliability, Readability, and Clinical Validity of ChatGPT 5.2

Sep 2026 · Archives of Current Medical Research · 0 citations · 23 references

TL;DR

ChatGPT 5.2 is a valuable AI–assisted chatbot that facilitates patient education by providing responses regarding sleeve gastrectomy that are generally accurate and acceptable, but the categorization of 4–16% of the responses as "Incorrect," the overall difficult readability levels, and the significant variability observed among evaluators with different levels of clinical experience underscore the indispensable role of clinical expertise and specialized professional guidance in surgical practice.

Abstract

Background: Although artificial intelligence (AI) has been used in patient education for some time, the accuracy, reliability, and clinical appropriateness of AI-generated medical content remain inadequately defined and continue to be debated. This study aimed to evaluate the reliability, readability, and comprehensibility of responses generated by ChatGPT 5.2 to the most frequently asked questions related to sleeve gastrectomy, and to assess the responses by surgeons with varying levels of clinical expertise. Methods: Twenty-four questions regarding sleeve gastrectomy were asked to ChatGPT 5.2, and responses were evaluated using a Likert-Scale by three general surgeons with varying levels of expertise. The readability and comprehensibility analysis was also conducted, using Flesch–Kincaid Grade Level and Flesch Reading Ease Scores. Results: Although excellent intra-rater reliability was observed, inter-rater reliability was poor. The evaluator with the greatest clinical expertise assigned the highest mean score (2.92±0.776), whereas the evaluator possessing predominantly theoretical knowledge assigned the lowest mean score (2.58±0.881). In the readability analysis, responses in all subcategories were classified as “difficult to read,” whereas only the responses within the postoperative course category were categorized as “fairly difficult”. Conclusions: ChatGPT 5.2 is a valuable AI–assisted chatbot that facilitates patient education by providing responses regarding sleeve gastrectomy that are generally accurate and acceptable. Nevertheless, the categorization of 4–16% of the responses as "Incorrect," the overall difficult readability levels, and the significant variability observed among evaluators with different levels of clinical experience underscore the indispensable role of clinical expertise and specialized professional guidance in surgical practice.

Read PDF

Similar papers

Open access Sep 2026

Accuracy and limitations of artificial intelligence chatbots in answering patient questions on scoliosis surgery.

Objective The aim of this study was to analyze the accuracy, reliability, and quality of the content of responses provided by artificial intelligence (AI)-based chatbots to frequently asked questions related to scoliosis surgery. Methods A set of 25 questions related to diagnosis, treatment options, surgical risks, a...

G. Alibakan, Y. Sulek · 0 citations
Review Open access Aug 2026

Quality, readability, and patient safety of ChatGPT-generated responses to fall-related questions in older adults: a multidisciplinary evaluation

While ChatGPT provided generally acceptable clinical information, variability in readability and expert ratings raises patient safety concerns and AI-generated health content should be reviewed and tailored to older adults' health literacy needs before clinical use.

Merve Arı, N. Ilçin, Hatice Yağcıoğlu et al. · 0 citations
#generative ai Open access Sep 2026

Generative AI responses to exercise-related questions asked by patients with multiple sclerosis: An evaluation of quality, accuracy, and readability.

ChatGPT-generated responses demonstrated higher quality, accuracy, and reliability than those generated by Gemini, Nonetheless, patient accessibility could be limited by the poor readability metrics observed in both tools.

A. A. Delibay, Cimen Olçay Demir, Nisa Turutgen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.