Evaluation of Small and Large Language Models for Calculation of the ASA Score and Charlson Comorbidity Index in Orthopedic Surgical Patients: A Retrospective Concordance Analysis
Aug 2026· Bioengineering· 0 citations· 31 references
TL;DR
Among the six evaluated model configurations, GPT-5.2 achieved significantly higher agreement with the clinician-derived composite reference than the other tested models for both ASA-PS and CCI in post hoc paired analyses with multiplicity correction.
Abstract
Background: The ASA Physical Status (ASA-PS) classification and the Charlson Comorbidity Index (CCI) are common pre-operative scoring tools. Language models could automate structured pre-operative scoring, but direct comparisons require paired inference because all models are evaluated on the same patients. Methods: In this retrospective single-center concordance analysis, 101 consecutive adult orthopedic patients were independently rated by two clinicians; the rounded mean for ASA-PS and arithmetic mean for CCI formed a clinician-derived composite reference. The cohort contained no ASA-PS IV-V patients. Six model configurations received identical prompts. Agreement was assessed using quadratic weighted kappa, ICC(2,1), exact and adjacent agreement, MAD, RMSE, and Bland–Altman limits. Post hoc between-model comparisons used 10,000 patient-level paired bootstrap replicates with Benjamini–Hochberg correction. Results: Inter-clinician weighted kappa was 0.713 for ASA-PS and 0.914 for CCI. GPT-5.2 reached kappa 0.884 for ASA-PS and 0.970 for CCI. In paired analyses, GPT-5.2 had significantly higher quadratic weighted kappa than every other tested model for both outcomes and significantly higher ICC for CCI after multiplicity correction. Phi4 and deepseek-r1-70B were not significantly different from inter-clinician agreement for CCI kappa or ICC; equivalence was not tested. Conclusions: Among the six evaluated model configurations, GPT-5.2 achieved significantly higher agreement with the clinician-derived composite reference than the other tested models for both ASA-PS and CCI in post hoc paired analyses with multiplicity correction. Locally deployable phi4 and deepseek-r1-70B showed CCI agreement estimates that were not statistically distinguishable from inter-clinician agreement, although equivalence was not tested. These findings are limited to the evaluated models and study cohort.
Background:Large language models (LLMs) are increasingly used by clinicians and trainees for perioperative decision support, yet their alignment with Enhanced Recovery After Surgery (ERAS) recommendations remains uncertain.Methods:We converted the 2025 ERAS Society recommendations for elective colorectal surgery into a 52‑question bank (preoperative (n=28), intraoperative (n=10), and postoperative (n=14). Each question was asked once, in a new chat, to Google Gemini and OpenAI ChatGPT (web interfaces; no follow‑up prompts or regeneration; queries performed on 17 Feb 2026). Responses were blinded as A/B and independently scored by two clinicians (an anesthesiologist and a general surgeon) for (i) guideline concordance on a 5‑point Likert scale and (ii) safety risk (0=none, 1=potential harm, 2=critical harm). Primary analysis used paired Wilcoxon signed‑rank tests (Likert) and exact McNemar tests (any safety flag ≥1). Inter‑rater agreement was estimated with quadratic weighted kappa (QWK).Results:A total of 208 ratings were generated (52 questions × 2 models × 2 raters). Mean Likert concordance was 4.49±0.62 for Gemini and 4.44±0.65 for ChatGPT (paired Wilcoxon p=0.604). Any safety flag occurred in 20.2% (Gemini) and 15.4% (ChatGPT) of ratings (McNemar p=0.383); no responses were rated as critical harm. Inter‑rater agreement was lower for Gemini (QWK=0.225) than for ChatGPT (QWK=0.597).Conclusions:Both LLMs showed high overall concordance with ERAS colorectal recommendations, with no significant overall difference in scores. However, safety‑relevant deviations were not rare, and rater agreement varied by model, highlighting the need for clinician oversight and standardized evaluation frameworks before bedside use.Keywords:ERAS; colorectal surgery; large language model; ChatGPT; Gemini;
Sezer Gökçen· Cukurova Anestezi ve Cerrahi...· 0 citations
Objective: To evaluate decision concordance between commercially available multimodal large language models (LLMs), resident doctors, and senior-surgeon ground truth for surgical indication and spinal level in degenerative lumbar spine disease. Methods: We retrospectively analyzed 147 consecutive patients. Each case included clinical documentation and MRI presented as two composite PNG images. Two resident doctors and three multimodal LLMs (GPT 5.5, Claude Sonnet 4.6, Gemini 3.1 Pro) independently assessed operative versus conservative management and, if operative, the surgical level. Analyses used Cochran's Q, McNemar tests with Holm correction, and Bayesian methods. Results: LLMs achieved higher therapy-decision accuracy (66.0%-68.0%; 97-100/147) than residents (54.4%; 80/147) but over-recommended surgery. Conditional level accuracy when surgery was correctly indicated was 71.4% (20/28) for residents versus 33.3%-41.1% for LLMs. Conclusion: Off-the-shelf multimodal LLMs approximate human performance for binary surgical indication but remain inferior for precise level localization. These results establish a practice-relevant baseline of spatial reasoning limitations for tools already used by patients and junior doctors.
M. Hamdan, A. Harati, A. Al-bakheet et al.· medRxiv· 0 citations
OBJECTIVE
The Risk Assessment and Prediction Tool (RAPT) has been utilized to anticipate discharge needs after procedures such as total joint arthroplasty. Its usefulness for spine patients, particularly those undergoing transforaminal lumbar interbody fusion (TLIF), has not been clearly established. This study evaluated the relationship between the preoperative RAPT score and 3 postoperative outcomes: discharge destination, hospital length of stay (LOS), and 30-day readmission.
METHODS
A retrospective cohort study was conducted of adults who underwent elective TLIF with a recorded preoperative RAPT score. RAPT was analyzed as a continuous variable. Home discharge and 30-day readmission were modeled with logistic regression, and LOS with linear regression. Multivariable models adjusted for age, sex, Charlson Comorbidity Index (CCI), and insurance type. Discrimination for facility discharge was assessed by receiver operating characteristic (ROC) analysis at literature-aligned thresholds (RAPT scores 9.5 and 8.5: balanced and conservative, respectively).
RESULTS
Among 116 patients, the mean age was 62.2 years and 50.9% of patients were female; the mean BMI was 29.9, and the mean CCI was 5.64. The mean RAPT score was 9.65, and the mean LOS was 3.14 days. Discharge to a skilled nursing or rehabilitation facility occurred in 6.9% of patients, and 30-day readmission occurred in 6.0%. Each 1-point increase in the RAPT score was associated with higher odds of home discharge (univariate: OR 1.74, 95% CI 1.11-2.73, p = 0.016; multivariable: OR 2.17, 95% CI 1.24-3.80, p = 0.007) and a shorter LOS (β = -0.38 days, 95% CI -0.71 to -0.05, p = 0.025; adjusted β = -0.39, bootstrap 95% CI -0.76 to -0.10, p = 0.038). The RAPT score was not associated with 30-day readmission (adjusted OR 0.93, 95% CI 0.52-1.65; p = 0.798). ROC analysis for predicting facility discharge showed moderate discrimination with an area under the curve of 0.709 (95% CI 0.543-0.876, p = 0.049), with sensitivity 63% and specificity 62% at 9.5, and sensitivity 38% and specificity 82% at 8.5. Youden's index revealed an optimal cutoff of 10.5, with sensitivity 100% and specificity 29.6%.
CONCLUSIONS
In patients undergoing TLIF, higher preoperative RAPT scores were associated with greater odds of home discharge and shorter LOS. RAPT may serve as a practical preoperative tool to support discharge planning and resource allocation in spine surgery.
Gabriel A. Gonzalez, Caden R. Moenning, Aaron Davidson et al.· Journal of Neurosurgery : Sp...· 0 citations
Objective
To evaluate the diagnostic performance of multiple large language models (LLMs) against expert consensus in determining surgical intervention needs for feline metacarpal and metatarsal fractures.
Methods
In this retrospective study (December 2023 to February 2025), 73 clinical cases of feline metacarpal and metatarsal fractures were evaluated. Two board-certified veterinary orthopedic surgeons established a reference standard for surgical versus conservative management. Five LLMs (ChatGPT, version 5.2 [OpenAI Inc]; Gemini, version 3 Pro [Alphabet Inc]; Grok, version 4.1 [SpaceXAI]; Qwen, version 3.5 [Alibaba Cloud]; and Claude Sonnet, version 4.5 [Anthropic PBC]; and Claude Sonnet, version 4.5 [Anthropic PBC]) assessed anonymized case summaries using a standardized zero-shot prompt. Model recommendations were compared with the reference standard to calculate accuracy, sensitivity, specificity, and the Cohen κ.
Results
The reference standard classified 49 cases (67.1%) as surgical and 24 (32.9%) as conservative. ChatGPT achieved the highest performance (accuracy, 84.9%; sensitivity, 79.6%; specificity, 95.8%; κ = 0.69). Other models showed lower performance; Qwen and Claude Sonnet failed to identify any surgical cases (0% sensitivity). A systematic bias toward conservative management was observed across all models, causing high false-negative rates (undertriage).
Conclusions
LLMs demonstrate highly variable diagnostic performance and a systematic risk of undertriage in feline fracture assessment. While top-performing models approach expert-level agreement, others fail in critical clinical scenarios.
Clinical Relevance
Independent LLM use for surgical decision-making in feline orthopedics is not recommended due to undertriage risks. These tools require strict clinician oversight for preliminary triage.
S. Okur, Ç. Özkalıpçı, Büşra Baykal et al.· Journal of the American Vete...· 0 citations
Non-structural clinical data, such as post-operative complications, are susceptible to inter-observer disagreement, undermining data validity. This study assesses the validity of multi-timepoint Clavien-Dindo complication grading in a single-procedure cohort of Whipple resections.
Complication grading within 7, 14, 30 and 90 days for 130 Whipple resections were independently coded by two final-year medical students and validated by a senior clinician. Inter-observer agreement was assessed using Cohen’s Kappa (κ). Disagreements were analysed by category (within minor, between major and minor, and within major grades). The senior clinician reviewed disagreements and identified potential systematic grading errors.
Across 1040 coding episodes (130 patients, two coders, four time intervals) agreement results showed: κ (days 1-7) = 0.77(95% CI, 0.66-0.88); κ (days 8-14) = 0.91(0.85-0.97); κ (days 15-30) = 0.88(0.81-0.95); κ (days 31-90) = 0.73(0.59-0.88). Recording the highest complication grade within 30 days reduced disagreements from 32 to 15, κ (days 1-30) = 0.84(0.50-1). Most disagreements occurred within minor grades (I–II). Disagreement between minor and major grades (≤II vs ≥IIIa) was only observed in days 1-7.
This study demonstrates that a supervised dual junior rater approach yields substantial-to-excellent inter-observer agreement in non-structural clinical data coding. However, agreement is likely over-estimated when only the highest grade is recorded over an extended interval (e.g. 30 days). For shorter time-series intervals, this approach is more prone to inconsistencies, highlighting the need for diagnostic criteria-based, automated data collection system to ensure data robustness for research and care quality improvement.
Sam Pathmanathan, Shi Lam, Sarah Alsaad et al.· British Journal of Surgery· 0 citations
OBJECTIVE
To determine whether presenting matched obstetric and gynecologic triage scenarios as patient-language prompts rather than clinician-language prompts affects expert-rated clinical confidence and the clinical safety of LLM-generated advice.
METHODS
Thirty obstetric and gynecologic scenarios were presented in matched clinician-language and patient-language Turkish formats to four LLMs. Five specialists independently evaluated 240 responses, generating 1,200 ratings. The primary outcome was the Global Clinical Confidence Score (GCCS; 0-2); five secondary outcomes were rated on 1-5 scales. Associations were examined using ordinal logistic generalized estimating equations adjusted for model and evaluator.
RESULTS
Clinically reliable responses (GCCS = 2) accounted for 89.3% of clinician-language and 91.3% of patient-language ratings. Patient-language phrasing was not significantly associated with overall GCCS (cumulative odds ratio 0.78, 95% confidence interval 0.57-1.06; p = 0.115), and the language-by-model interaction was not significant (p = 0.422). Patient-language prompts were associated with fewer GCCS = 0 ratings in a binary generalized estimating equations analysis (odds ratio 0.66, 95% confidence interval 0.46-0.96; p = 0.031), although the exact paired McNemar test was not significant (p = 0.096). After false-discovery-rate correction, patient-language prompts had higher evaluator-level triage appropriateness and clinical applicability scores (both adjusted p = 0.028).
CONCLUSION
No significant difference in overall expert-rated clinical confidence was detected between patient-language and clinician-language prompts.
Onur Ada, Uğurcan Dağlı, E. Bilen et al.· European Journal of Obstetri...· 0 citations
A new method for surgically removing training examples from a model reveals that as datasets grow, the link between what a model learns and what it produces dissolves.