Back to feed
Review Open access

184 Evaluating Reasoning-Tuned Large Language Models for Clinical Decision-Making in Spine Surgery

Jul 2026 · British Journal of Surgery · 0 citations

Abstract

Large language models (LLMs) such as OpenAI o1 and DeepSeek-R1 are designed to move beyond factual recall by modelling deliberative thought processes. Their performance in complex spine scenarios remains unclear. This study evaluates whether reasoning-tuned LLMs can generate coherent and clinically relevant management plans when assessed by fellowship-trained spine surgeons. Eleven synthetic case vignettes were presented to OpenAI o1 (full) and DeepSeek R1 with identical prompts requesting diagnostic impression, reasoning, and management plan. Outputs were anonymised and randomised for blind review. Eight fellowship-trained spine surgeons [five consultants, three fellows] from the United Kingdom, Switzerland, Nigeria, and Zambia scored diagnostic accuracy, reasoning, surgical plan appropriateness, and clarity on five-point Likert scales. Eighty-six paired evaluations were analysed using two-tailed paired t-tests with Bonferroni correction, adjusted α=0.0125. OpenAI o1 (full) outperformed DeepSeek R1 across all domains. Means [Standard Deviation] and p values were diagnostic accuracy 4.57 [0.60] vs 4.31 [0.79], p < 0.001, reasoning and thoroughness 4.48 [0.68] vs 4.23 [0.75], p = 0.005, surgical plan appropriateness 4.33 [0.76] vs 4.06 [0.86], p = 0.010, clarity 4.47 [0.68] vs 4.15 [0.85], p < 0.001. All comparisons met the corrected significance threshold, and o1 showed lower standard deviations, which signals more consistent quality across raters and cases. Reasoning-tuned LLMs can emulate elements of expert surgical decision-making. OpenAI o1 (full) produced more accurate, thorough, appropriate, and clear plans, with greater consistency, while DeepSeek R1 showed credible but more variable outputs. Transparent validation and reporting remain essential before clinical use.

Read PDF