Author

A. Ravishankar

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Review Open access Jul 2026

184 Evaluating Reasoning-Tuned Large Language Models for Clinical Decision-Making in Spine Surgery

Large language models (LLMs) such as OpenAI o1 and DeepSeek-R1 are designed to move beyond factual recall by modelling deliberative thought processes. Their performance in complex spine scenarios remains unclear. This study evaluates whether reasoning-tuned LLMs can generate coherent and clinically relevant management plans when assessed by fellowship-trained spine surgeons. Eleven synthetic case vignettes were presented to OpenAI o1 (full) and DeepSeek R1 with identical prompts requesting diagnostic impression, reasoning, and management plan. Outputs were anonymised and randomised for blind review. Eight fellowship-trained spine surgeons [five consultants, three fellows] from the United Kingdom, Switzerland, Nigeria, and Zambia scored diagnostic accuracy, reasoning, surgical plan appropriateness, and clarity on five-point Likert scales. Eighty-six paired evaluations were analysed using two-tailed paired t-tests with Bonferroni correction, adjusted α=0.0125. OpenAI o1 (full) outperformed DeepSeek R1 across all domains. Means [Standard Deviation] and p values were diagnostic accuracy 4.57 [0.60] vs 4.31 [0.79], p < 0.001, reasoning and thoroughness 4.48 [0.68] vs 4.23 [0.75], p = 0.005, surgical plan appropriateness 4.33 [0.76] vs 4.06 [0.86], p = 0.010, clarity 4.47 [0.68] vs 4.15 [0.85], p < 0.001. All comparisons met the corrected significance threshold, and o1 showed lower standard deviations, which signals more consistent quality across raters and cases. Reasoning-tuned LLMs can emulate elements of expert surgical decision-making. OpenAI o1 (full) produced more accurate, thorough, appropriate, and clear plans, with greater consistency, while DeepSeek R1 showed credible but more variable outputs. Transparent validation and reporting remain essential before clinical use.

A. Ravishankar, C. Lam, A. Bulloso et al. · 0 citations