Aug 2026· European spine journal· 0 citations· 13 references
Medicine
TL;DR
LLMs can support communication and education following spine surgery when used with structured prompting when used with structured prompting and ChatGPT and Claude showed the highest correctness and completeness, particularly for practitioner-directed answers.
Abstract
Purpose
To evaluate whether large language models (LLMs) can provide accurate, complete, and audience-adapted answers to common spine-surgery-related questions for patients and family practitioners.
Methods
Ten frequently asked spine-surgery questions were collected at a level 1 trauma center and simplified linguistically. Five LLMs (ChatGPT, Claude 3.5 Sonnet, Gemini Advanced 1.5 Pro, Copilot Pro, and DeepSeek V3) were queried using zero-shot prompting with persona-specific instructions for family practitioners and middle-aged patients. Responses were assessed by spine surgeons and non-medical raters for correctness, completeness, adaptability, and empathy using five-point Likert scales. Readability was quantified using the Flesch Reading Ease Score (FRES).
Results
All LLMs generated largely correct and usable responses. ChatGPT and Claude showed the highest correctness and completeness, particularly for practitioner-directed answers. Gemini and Copilot achieved superior readability and empathy for patient-facing responses. DeepSeek demonstrated balanced performance across all domains. Readability differed substantially between practitioner- and patient-oriented outputs.
Conclusion
LLMs can support communication and education following spine surgery when used with structured prompting. Clinical oversight remains essential to mitigate risks related to inaccuracies and hallucinations.
LEVEL OF EVIDENCE
III.
INTRODUCTION
Large language models (LLMs) are increasingly used as sources of information across many fields, including healthcare. As patients turn to these models for health-related queries, evaluating the accuracy and reliability of their responses is essential. Peri-acetabular osteotomy (PAO) is performed on younger patients - a group more likely to use digital tools like LLMs for health information. This study assesses the accuracy and readability performance of two leading LLMs, ChatGPT and Google Gemini, in addressing common patient questions on PAO.
METHODS
A panel of fellowship-trained PAO surgeons created ten commonly asked patient questions based on real-world experience. Responses from each LLM were assessed by the same three surgeons, blinded to response origin, using a 5-point Likert scale to evaluate clarity, accuracy, and completeness. Readability was measured with Flesch-Kincaid Reading scores.
RESULTS
ChatGPT outperformed Gemini with an average score of 4.17 vs 3.13 (t = -3.08, p = 0.006). ChatGPT's responses were often rated higher for completeness and clarity, particularly in areas needing detailed explanation, and usually required minimal clarification. Gemini sometimes lacked specificity or included minor inaccuracies that reduced its perceived reliability. Both LLMs produced responses with similar "difficult" Flesch Reading Ease scores.
CONCLUSIONS
There may be significant differences in how effectively LLMs support patients with surgical queries. ChatGPT more consistently met expert standards for clarity and thoroughness. As LLM usage expands, ChatGPT may aid patient education on hip surgery, supporting consultations, informed decisions and postoperative guidance.
TP Davis, B. Guevel, K. Logishetty et al.· Annals of the Royal College...· 0 citations
PURPOSE
To evaluate and compare the performance of four general-purpose large language models (LLMs) (ChatGPT-5, Claude 4, Grok 4 and Gemini 2.5) in answering specialised clinical questions related to total knee arthroplasty (TKA) derived from the World Expert Meeting in Arthroplasty (WEMA).
METHODS
This is a cross-sectional comparative study. Twenty questions on TKA supported by moderate-strong level of evidence were randomly selected from the WEMA. Three orthopaedic surgeons independently performed a blinded assessment of all LLM-generated responses. An adapted version of the QUEST rating system, a comprehensive framework designed for the objective human assessment of LLM performance across healthcare-related subdomains, was used. Furthermore, the same three evaluators subjectively selected the best-performing LLM response for each question.
RESULTS
The four LLMs presented statistically significant differences in overall performance based on the QUEST framework (score range 1-5): Gemini 2.5; 4.77 ± 0.06, Claude 4; 4.72 ± 0.07, ChatGPT-5; 4.70 ± 0.08 and GROK 4; 4.62 ± 0.09 (p < 0.001). Gemini 2.5 achieved the highest scores in the Accuracy (4.58 ± 0.70), Comprehensiveness (4.87 ± 0.34) and Trust (4.52 ± 0.62) dimensions. However, Claude 4 obtained the highest score for the Currency (4.23 ± 0.67) dimension. When assessors subjectively selected the superior answer for each question, Claude 4 was chosen most frequently, in 46.7% of cases, followed by ChatGPT-5 in 25.4%, Gemini 2.5 in 22.9% and GROK 4 in 7.5% of cases.
CONCLUSIONS
Four general-purpose LLMs (ChatGPT-5, Claude 4, Grok 4 and Gemini 2.5) demonstrated good overall performance when addressing specialised clinical questions related to TKA. No single model consistently outperformed the others across all evaluated domains.
LEVEL OF EVIDENCE
Level V.
Oriol Pujol, R. Ferrer, Alex Coelho et al.· Knee Surgery, Sports Traumat...· 0 citations
Background Patients are increasingly turning to online resources and artificial intelligence (AI)-based tools to obtain information about orthopedic conditions and surgical options. Large language models, such as ChatGPT, are becoming prominent in patient education; however, their reliability and readability remain uncertain. This study evaluated the quality and readability of responses generated by ChatGPT-4o and ChatGPT-5 to frequently asked patient questions regarding hallux rigidus fusion surgery. Methods Twenty commonly asked patient questions were compiled and presented to ChatGPT-4o and ChatGPT-5. Readability was assessed using the Flesch–Kincaid Grade Level, Gunning Fog, Coleman–Liau, and Simple Measure of Gobbledygook indices. Quality was evaluated with the DISCERN tool, response accuracy scores, and Journal of the American Medical Association (JAMA) criteria. Interrater agreement was measured using the Intraclass Correlation Coefficient (ICC). Results ChatGPT-4o generated longer responses (802 vs. 242 words; p<0.001) with slightly higher readability grade levels (10.81 vs. 10.37; p=0.031). Accuracy (2.00 vs. 1.85; p=0.323) and DISCERN scores (49.35 vs. 48.93; p=0.747) showed no significant differences. All responses received a JAMA score of 0 due to the absence of citations, authorship, or transparency indicators. Interrater reliability indicated moderate to good agreement (ICC: 0.68–0.80). Conclusion ChatGPT-4o and ChatGPT-5 provide generally satisfactory yet non-comprehensive, limited-quality information at a level above tenth-grade regarding hallux rigidus fusion surgery. Although linguistically coherent, responses lack evidence-based detail and individualized guidance. These models may supplement, but cannot replace, expert orthopedic counseling. Ensuring physician oversight and integrating validated, updated clinical content remain essential for safe implementation of AI-generated patient information.
Kamil Balaban, Mehmet Batu Ertan, Mahmut Kalem· Digital Health· 0 citations
BACKGROUND
Patients increasingly turn to AI chatbots for medical information, including before complex orthopaedic procedures such as progressive collapsing foot deformity (PCFD) surgery. Whether these tools deliver content of sufficient quality and accessibility for preoperative patient education remains unclear, particularly across competing platforms. This study addressed three questions: (1) Do ChatGPT, Perplexity AI and Google Gemini differ in the accuracy, comprehensiveness and clarity of their responses to PCFD-related patient questions? (2) Do these platforms produce content meeting recommended readability thresholds for patient education? (3) Does the level of agreement among blinded foot and ankle surgeons rating the quality of chatbot responses vary depending on the platform used?
HYPOTHESIS
The three AI chatbot platforms produce responses of comparable accuracy but differ significantly in readability, with none reaching the recommended readability thresholds for patient education materials.
PATIENTS AND METHODS
Cross-sectional comparative study. Twenty frequently asked questions regarding PCFD, covering disease understanding, conservative management, surgical planning and postoperative recovery, were submitted verbatim to ChatGPT (GPT-4o mini), Perplexity AI and Google Gemini (free versions, March 25, 2026). The 60 resulting responses were rated by three blinded foot and ankle surgeons on three 5-point Likert scales (accuracy, comprehensiveness, clarity). Readability was assessed using the Flesch Reading Ease (FRE) and Flesch-Kincaid Grade Level (FKGL). Inter-rater agreement used Kendall's W; differences between platforms were analysed using Kruskal-Wallis tests with Bonferroni-corrected pairwise comparisons.
RESULTS
All platforms produced responses rated accurate to very accurate. Perplexity achieved significantly higher accuracy than ChatGPT (p = 0.0003) and higher accuracy and clarity than Gemini (p = 0.0065 and p = 0.0075). No platform reached the recommended FRE ≥ 60 or FKGL ≤ 6 thresholds: median FKGL ranged from 13.1 (ChatGPT) to 21.4 (Perplexity), with ChatGPT producing the most readable and Perplexity the least readable content (p < 0.001). Inter-rater agreement was fair to substantial across platforms, lowest for Perplexity.
DISCUSSION
AI chatbots produce generally accurate baseline information on PCFD surgery, with Perplexity showing significantly higher expert-rated accuracy and clarity than the other platforms-contrary to our hypothesis of comparable accuracy-while readability remains uniformly inadequate for all platforms, as hypothesized. These tools may serve as a supplementary source of information, but their inadequate readability suggests they are not yet suited to replace tailored, surgeon-led patient education.
LEVEL OF EVIDENCE
III; cross-sectional comparative study.
L. Micicoi, J. Brué, Saumith Menon et al.· Orthopaedics & Traumatology:...· 0 citations
Aims: Large language models (LLMs) have demonstrated promising performance in medical examinations; however, their effectiveness in orthopedics-particularly in visually demanding questions-remains unclear. This study aimed to compare the performance of ChatGPT-5.2 and Gemini 3 Pro on orthopedics and traumatology questions from the Turkish Medical Specialization Examination (TUS).Methods: A total of 7.200 TUS questions from 2006 to 2021 were screened, and 163 orthopedics-related questions with validated answer keys were included. Questions were categorized into seven subspecialties and classified as text-based (n=154) or visual (n=9). Each question was independently presented to both models in Turkish without prior prompting. Responses were recorded as correct or incorrect. Statistical analyses included the McNemar test for paired comparisons and Fisher’s exact test for within-model comparisons of visual versus text-based performance.Results: ChatGPT-5.2 achieved an accuracy of 93.3% (152/163), while Gemini 3 Pro achieved 92.0% (150/163), with no statistically significant difference between models (p=0.593). Both models showed higher error rates on visual questions than on text-based questions (ChatGPT-5.2: 44.4% vs. 4.5%, p=0.001; Gemini 3 Pro: 33.3% vs. 6.5%, p=0.025); however, these visual comparisons were based on only nine items and should be regarded as exploratory. Error distribution across subspecialties was similar between models, with orthopedic trauma showing the highest error frequency.Conclusion: Both ChatGPT-5.2 and Gemini 3 Pro demonstrated high overall accuracy in orthopedic examination questions, with no significant difference between models. However, both models exhibited a marked decline in performance on visually based questions, highlighting persistent limitations in multimodal reasoning. Future research should incorporate larger visual datasets to better evaluate AI performance in orthopedic contexts.
Batuhan Ayhan, Samet Batuhan Yoğurt· Anatolian Current Medical Jo...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.