Skip to content
Open access

Multimodal Large Language Models vs. Medical Doctors in Degenerative Lumbar Spine Surgery: A Retrospective Decision Concordance Study of 147 Patients

Aug 2026 · medRxiv · 0 citations
Medicine

TL;DR

Off-the-shelf multimodal LLMs approximate human performance for binary surgical indication but remain inferior for precise level localization, establish a practice-relevant baseline of spatial reasoning limitations for tools already used by patients and junior doctors.

Abstract

Objective: To evaluate decision concordance between commercially available multimodal large language models (LLMs), resident doctors, and senior-surgeon ground truth for surgical indication and spinal level in degenerative lumbar spine disease. Methods: We retrospectively analyzed 147 consecutive patients. Each case included clinical documentation and MRI presented as two composite PNG images. Two resident doctors and three multimodal LLMs (GPT 5.5, Claude Sonnet 4.6, Gemini 3.1 Pro) independently assessed operative versus conservative management and, if operative, the surgical level. Analyses used Cochran's Q, McNemar tests with Holm correction, and Bayesian methods. Results: LLMs achieved higher therapy-decision accuracy (66.0%-68.0%; 97-100/147) than residents (54.4%; 80/147) but over-recommended surgery. Conditional level accuracy when surgery was correctly indicated was 71.4% (20/28) for residents versus 33.3%-41.1% for LLMs. Conclusion: Off-the-shelf multimodal LLMs approximate human performance for binary surgical indication but remain inferior for precise level localization. These results establish a practice-relevant baseline of spatial reasoning limitations for tools already used by patients and junior doctors.

Read PDF

Similar papers

Review Open access Jul 2026

Large Language Models in Spine Surgery: A Clinical Decision- Making Framework for the Next Decade with Emphasis on Degenerative Spine Care and LMIC Applications

Degenerative spine disorders are a leading cause of disability worldwide and impose a growing burden on health systems, particularly in low- and middle-income countries. Large language models are emerging as powerful tools capable of supporting clinical decision-making, synthesising complex evidence, generating patient-specific explanations, and improving clinical documentation. Their ability to process free-text clinical narratives and integrate multiple sources of information makes them particularly suited to degenerative spine care, where decision-making requires the integration of symptoms, neurological findings, imaging, and patient preferences. This narrative review examines the evolving role of artificial intelligence in spine surgery, with particular emphasis on large language models in degenerative conditions of the cervical and lumbar spine. We describe current and emerging applications across clinical triage, radiological interpretation, guideline synthesis, patient communication, and workflow optimisation. Building on these insights, we propose a structured six-level clinical decision-making framework spanning initial patient contact to postoperative care. We also discuss key ethical, medico-legal, and governance considerations relevant to the safe implementation of these technologies, particularly in resource-constrained environments. Large language models are unlikely to replace clinical judgement; however, when integrated within structured workflows and appropriate safety systems, they have the potential to enhance the quality, efficiency, and equity of degenerative spine care globally over the coming decade.

S. Ganesh, Jeena Joseph · 0 citations
Open access Aug 2026

Prompt Configurations for Multimodal Large Language Models in Diagnosing and Staging Osteonecrosis of the Femoral Head: Multimodel Retrospective Observational Diagnostic Study

Abstract Background Multimodal large language models (MLLMs) have emerging potential for interpreting medical images and text, but their performance in orthopedic imaging tasks and the influence of prompt configuration remain insufficiently studied. Objective This study aimed to evaluate the performance of commercial and open-source MLLMs for diagnosing and staging osteonecrosis of the femoral head (ONFH) and to assess how different prompt configurations affect model performance. Methods This single-center retrospective diagnostic accuracy study included 159 radiograph patients contributing 318 hip-level observations and 170 magnetic resonance imaging (MRI) patients contributing 340 hip-level observations; 55 patients with both modalities formed the multi-image (MI) subgroup between July 2023 and December 2024. Four MLLMs were evaluated: GPT-4o, Claude 3.7 Sonnet, Qwen2.5-VL 72B, and Gemma 3 27B. Three prompt configurations were tested: single image (SI), image plus radiology description (ID), and MI. Model performance was assessed for ONFH detection; early- versus late-stage differentiation; detailed grading using the Ficat, Association Research Circulation Osseous (ARCO), and Steinberg systems; and grading reliability using intraclass correlation coefficients (ICCs). Results Model performance varied by prompt configuration and imaging input. For ONFH detection, the SI configuration yielded a mean detection area under the receiver operating characteristic curve (AUC) of 0.55 (SD 0.03), whereas the ID configuration achieved a mean detection AUC of 0.91 (SD 0.01). In radiograph-based ONFH detection, ID input achieved a mean accuracy of 0.88 (SD 0.01); in MRI-based ONFH detection, ID input achieved a mean accuracy of 0.85 (SD 0.01). For early- versus late-stage differentiation, the mean accuracy was 0.65 (SD 0.11) with SI input, 0.78 (SD 0.04) with ID input, and approximately 0.59 (SD 0.10) with MI input. For detailed grading, ID input improved mean accuracy across the Ficat, ARCO, and Steinberg systems compared with SI input. In ARCO grading reliability analysis, the mean MLLM ICC was 0.51 (SD 0.18) for SI, 0.97 (SD 0.02) for ID, and 0.49 (SD 0.15) for MI; surgeon interrater and intrarater ICCs were 0.70 and 0.81, respectively. In the commercial versus open-source model comparison, no significant overall difference was observed between model groups (P=.83). Conclusions Prompt configuration strongly influenced MLLM performance in ONFH diagnosis and staging. Pairing images with deidentified radiology descriptions improved diagnostic and grading performance, whereas MI input did not provide consistent additional benefit in this retrospective single-center evaluation. These findings support the potential role of MLLMs as assistive tools in human-AI orthopedic imaging workflows, but external validation, careful input standardization, and prospective clinical evaluation are needed before clinical deployment.

Jiesheng Zhu, Xingxing Huang, Jincheng Shi et al. · 0 citations
Review Aug 2026

Large language model treatment-pathway outputs based on structured clinical text in mid- and low rectal cancer: Concordance with multidisciplinary team decisions and features associated with discordance.

INTRODUCTION This study evaluated concordance between treatment-pathway outputs generated from structured clinical text by the large language model (LLM) GPT-5.4 Thinking and multidisciplinary team (MDT) decisions for mid- and low rectal cancer. MATERIALS AND METHODS This single-center retrospective study included 260 patients who underwent standardized assessment, MDT discussion, and curative-intent surgery between January 2022 and December 2023. After database lock, structured de-identified clinical text was entered into GPT-5.4 Thinking using prespecified preoperative and postoperative templates, with MDT decisions treated as real-world reference decisions. The primary endpoint was preoperative concordance; secondary endpoints included postoperative and overall concordance. Concordance metrics, directional discordance, baseline comparators, repeatability in a 50-case subset, and exploratory logistic regression models were assessed. RESULTS Preoperative, postoperative, and overall concordance rates were 57.7%, 66.9%, and 46.9%, respectively; Cohen's κ values were 0.269 and 0.449 for the preoperative and postoperative stages. Preoperatively, directional discordance relative to MDT decisions included 48 potential under-intensification, 28 potential over-intensification, and 34 directionally indeterminate or heterogeneous cases. Compared with the majority-class baseline, the LLM had the same preoperative crude concordance but higher κ and balanced category-specific concordance; postoperatively, it exceeded the majority-class baseline across these metrics. Non-identical mapped categories across three runs occurred in 12/50 preoperative and 8/50 postoperative assessments despite identical inputs. CONCLUSION GPT-5.4 Thinking showed some concordance with MDT decisions, but strict two-stage overall concordance remained limited. Large language models should be regarded as adjunctive, reviewable decision-support tools rather than replacements for MDTs.

Yuanze Wei, Yulong Tian, Xiaodong Liu et al. · 0 citations
#small language model Open access Aug 2026

Evaluation of Small and Large Language Models for Calculation of the ASA Score and Charlson Comorbidity Index in Orthopedic Surgical Patients: A Retrospective Concordance Analysis

Among the six evaluated model configurations, GPT-5.2 achieved significantly higher agreement with the clinician-derived composite reference than the other tested models for both ASA-PS and CCI in post hoc paired analyses with multiplicity correction.

Marco di Maio, G. Stopper, Vincenzo Di Matteo et al. · 0 citations
Review Open access Jul 2026

184 Evaluating Reasoning-Tuned Large Language Models for Clinical Decision-Making in Spine Surgery

Large language models (LLMs) such as OpenAI o1 and DeepSeek-R1 are designed to move beyond factual recall by modelling deliberative thought processes. Their performance in complex spine scenarios remains unclear. This study evaluates whether reasoning-tuned LLMs can generate coherent and clinically relevant management plans when assessed by fellowship-trained spine surgeons. Eleven synthetic case vignettes were presented to OpenAI o1 (full) and DeepSeek R1 with identical prompts requesting diagnostic impression, reasoning, and management plan. Outputs were anonymised and randomised for blind review. Eight fellowship-trained spine surgeons [five consultants, three fellows] from the United Kingdom, Switzerland, Nigeria, and Zambia scored diagnostic accuracy, reasoning, surgical plan appropriateness, and clarity on five-point Likert scales. Eighty-six paired evaluations were analysed using two-tailed paired t-tests with Bonferroni correction, adjusted α=0.0125. OpenAI o1 (full) outperformed DeepSeek R1 across all domains. Means [Standard Deviation] and p values were diagnostic accuracy 4.57 [0.60] vs 4.31 [0.79], p < 0.001, reasoning and thoroughness 4.48 [0.68] vs 4.23 [0.75], p = 0.005, surgical plan appropriateness 4.33 [0.76] vs 4.06 [0.86], p = 0.010, clarity 4.47 [0.68] vs 4.15 [0.85], p < 0.001. All comparisons met the corrected significance threshold, and o1 showed lower standard deviations, which signals more consistent quality across raters and cases. Reasoning-tuned LLMs can emulate elements of expert surgical decision-making. OpenAI o1 (full) produced more accurate, thorough, appropriate, and clear plans, with greater consistency, while DeepSeek R1 showed credible but more variable outputs. Transparent validation and reporting remain essential before clinical use.

A. Ravishankar, C. Lam, A. Bulloso et al. · 0 citations