Skip to content
Review Open access

Large language models in spine care and research.

Aug 2026 · European spine journal · 0 citations · 31 references
Medicine

TL;DR

This review highlights the promise of LLMs in enhancing decision support, workflow efficiency, research productivity, and patient communication in spine care, while emphasizing the need for interdisciplinary collaboration, robust evaluation metrics, and governance frameworks that prioritize patient safety and equity.

Abstract

Background

Large language models (LLMs) have emerged as powerful transformer-based systems capable of capturing long-range dependencies and complex semantic relationships in clinical language. In this review paper, we first examine the technical foundations of medical LLMs, including transformer architecture, attention mechanisms, training paradigms, and retrieval-augmented generation.

Results

We then survey documented applications in spine surgery and spinal care, highlighting moderate guideline concordance (46-67%) for diagnostic support, automated generation of operative notes and discharge summaries for administrative workflows, LLM-assisted literature review and manuscript drafting for research support (with ~68% novelty accuracy), and translation of complex surgical concepts into patient-friendly materials at a seventh-grade reading level. We next explore emerging multimodal models that integrate text, imaging, laboratory, and genomic data via cross-modal attention, demonstrating superior performance in holistic diagnostic and prognostic tasks.

Discussion

Finally, we discuss key implementation challenges, including model accuracy and hallucinations; computational, privacy, and regulatory constraints under HIPAA/GDPR; and bias mitigation, to outline strategies for safe, effective, and equitable deployment.

Conclusion

By mapping technical capabilities to clinical and research use cases, this review highlights the promise of LLMs in enhancing decision support, workflow efficiency, research productivity, and patient communication in spine care, while emphasizing the need for interdisciplinary collaboration, robust evaluation metrics, and governance frameworks that prioritize patient safety and equity.

Read PDF

Similar papers

Review Open access Sep 2026

Modern generative large language models in epilepsy care: a scoping review of current applications, challenges, and future directions

Modern generative large language models (LLMs) are increasingly being evaluated in epilepsy-related clinical tasks, but the evidence remains fragmented and their safe clinical role is uncertain. We conducted a scoping review following the PRISMA-ScR framework, searching PubMed, Embase, and the Web of Science Core Collection from inception to May 2026, to map current applications of modern generative LLMs in epilepsy-related clinical contexts and identify their reported benefits, limitations, and research gaps. Eligible studies evaluated generative LLMs or generative foundation models as the primary analytical or assistive engine in clinical or clinically oriented epilepsy tasks; non-generative natural language processing, conventional machine-learning, and signal-modeling studies served as contextual comparators only. Two reviewers independently screened records and extracted data for descriptive mapping and narrative synthesis. Twenty-four studies were included. Evidence was relatively more developed for text-centered tasks, including information extraction from electronic health records, question answering, and documentation support, while applications in differential diagnosis, presurgical evaluation, prognostic assessment, and patient communication remained early-stage. Across studies, the evidence base was heterogeneous, with frequent reliance on retrospective or simulated designs, single-center data, and limited external validation, alongside persistent concerns about hallucination, bias, model opacity, and workflow integration. Current evidence suggests a limited role for cautious, clinician-supervised use of generative LLMs in selected text-heavy epilepsy tasks, particularly extraction, summarization, documentation support, and patient education, but falls short of justifying independent or routine clinical deployment. Future studies should use narrower task definitions, epilepsy-specific datasets, transparent reporting of model versions and prompt design, external validation, and explicit documentation of clinician verification workflows.

Shi-Hao Ge, Yue-Qian Sun, Qun Wang · 0 citations
#generative ai Review Open access Sep 2026

Large language models in the vertical integration of spine surgery workflow: a scoping review.

PURPOSE To evaluate current evidence regarding the clinical reliability and reasoning capabilities of Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) within the spine surgery workflow. This scoping review utilizes a novel Vertical Integration Maturity Scale (VIMS) to map model maturity across five stages of the vertical workflow, identifying persistent research gaps and technical prerequisites for clinical implementation. METHODS A systematic search was conducted across PubMed, Embase, and Scopus for peer-reviewed studies published between January 2023 and January 2026. Utilizing the Population, Concept, and Context (PCC) framework, studies were selected based on their application of generative AI to the clinical evaluation and surgical management of spinal pathologies. Data were analyzed via independent dual-coding using a two-dimensional framework mapping five VIMS levels across five functional workflow stages. Synthesis included categorization by model architecture, input modality, and performance benchmarks, with inter-rater reliability calculated to validate the novel framework. RESULTS Of 351 identified records, 40 studies met the inclusion criteria. Research density peaked in Stage III (Evaluative Decision Logic) at VIMS Level 1 (n = 15), indicating a methodological focus on evidence-based guideline retrieval over patient-specific synthesis. While ChatGPT-4 showed high concordance (61.1% to 88.2%) with clinical guidelines, diagnostic precision fell to 11.1% in complex spinal deformity. A critical evidence gap persists in Stage IV (Tactile Procedure Planning) with a total absence of studies evaluating autonomous workflow agency (VIMS 4-5). This absence likely reflects a limitation of retrospective study designs reliant on proprietary, non-version-controlled models. While qualitative synthesis suggests hybrid architectures and Retrieval-Augmented Generation (RAG) may bridge existing semantic gaps, the inferential strength of these trends is currently limited by heterogenous benchmarking metrics. CONCLUSION Current LLMs demonstrate proficiency in foundational medical knowledge, but a significant evidence gap remains regarding their inferential validity in the non-deterministic surgical scenarios. Advancing clinical integration requires modular multimodal ensembles that process radiographic data alongside domain-specific reasoning scaffolds. Crucially, future research must transition from retrospective algorithmic benchmarking toward standardized, prospective clinical trials to evaluate algorithmic decision-making against real-world surgical outcomes.

Edwin H. Y. Lui, Ralph J. Mobbs · 0 citations
Review Open access Jul 2026

Transformer-Based Language Models for Clinical Decision Support Using Clinical Notes: A Scoping Review

Transformer-based language models have been applied across diverse clinical-note tasks, but the evidence base more strongly supports retrospective task feasibility than transportability, equitable performance, workflow benefit, or safe clinical deployment.

Saahoon Hong, Hunhui Na · 0 citations
#large language models Review Open access Sep 2026

Generative large language models in medicine: a scoping review of recent methodological advances

Generative large language models (LLMs) are rapidly transforming medicine, demonstrating unprecedented capability across a broad spectrum of clinical and biomedical tasks. While prior literature has extensively investigated their applications, the methodological foundation underpinning these models remains comparatively underexamined. In this review, we provide a mechanically grounded analysis of recent methodological advances shaping the development and deployment of generative LLMs in healthcare. We categorize the technical landscape into three principal pillars: pretraining, fine-tuning, and prompt engineering, and examine their key architectures, subtypes, and adaptation strategies based on literature published between 2023 and 2025. We further discuss the emerging directions, including efficient model infrastructures and LLMs-powered multi-agent systems, alongside critical challenges related to bias, generalization, and evaluation. By tracing the evolutionary trajectories of these methodologies, this scoping review provides a mechanism-centered framework to inform responsible model development and deployment in medical settings, tailored to task complexity, data characteristics, and resource constraints.

Fang Li, Jianfu Li, Weiguo Cao et al. · 0 citations
Review Open access Jul 2026

Aligning Clinical Needs and AI Capabilities: A Survey on LLMs for Medical Reasoning

A dual-view approach that connects clinical practice with computational methods is presented, establishing a five-level competency scheme following Miller’s Pyramid and linking deductive, inductive, and abductive reasoning patterns to common medical goals and tasks.

Qi Peng, Jiatong Li, Sirui Huang et al. · 6 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.