OrthoPilot is a clinical artificial intelligence system powered by a large language model (LLM) that integrates hospital data streams with authoritative external knowledge for continuous musculoskeletal care and autonomously retrieves real-time imaging, laboratory, pathology and order data and translates evolving patient states into evidence-based decisions from admission diagnosis through rehabilitation planning.
Abstract
Musculoskeletal diseases are among the leading causes of disability and drive the greatest global need for rehabilitation. Because recovery, remodelling and degeneration of bones, joints and related tissues unfold over months to years, care requires longitudinal management rather than isolated decisions. Clinicians must repeatedly integrate evolving patient evidence, medical knowledge and stage-specific functional goals, yet evidence is often fragmented across visits, departments and hospital systems, disrupting continuous, individualised management. Here we report OrthoPilot, a clinical artificial intelligence (AI) system powered by a large language model (LLM) that integrates hospital data streams with authoritative external knowledge for continuous musculoskeletal care. It autonomously retrieves real-time imaging, laboratory, pathology and order data and translates evolving patient states into evidence-based decisions from admission diagnosis through rehabilitation planning. We established a specialist-validated benchmark from real-world electronic health records (EHRs) spanning 1,000 disease codes. In a full-pathway reader study against 81 orthopaedic physicians, OrthoPilot outperformed experts with 25 years of experience in diagnostic reasoning, clinical decision-making and management planning. This advantage generalised across 60 external clinical centres, where OrthoPilot surpassed all evaluated intelligent systems. In a prospective physician decision-making study of 1,870 complex cases, OrthoPilot improved full-chain management success by 10.6%. In a randomised deployment involving 8,240 inpatients, integration into routine care increased cumulative cases per bed by 9.7% and improved patient-reported access to health information. These results move clinical AI from predicting isolated events toward executing longitudinal management across complete musculoskeletal care pathways.
Distal radius fractures are among the most common fragility injuries in older adults and account for a substantial share of orthopaedic trauma workload, particularly in high-volume and resource-constrained settings. In elderly patients the choice between nonoperative management and surgical fixation is rarely straightforward: long-term functional outcomes tend to converge across treatments, whereas the optimal approach for an individual is shaped by fracture pattern, bone quality, frailty, cognition, social circumstances and rehabilitation access. This narrative review, written for resource-constrained health systems and using Türkiye as a worked example, synthesises current evidence on conservative versus surgical treatment of distal radius fractures in older adults and examines where artificial intelligence (AI) and machine learning (ML) might realistically contribute along the care pathway. Sources were identified through PubMed and Web of Science searches of English-language literature conducted in 2025, prioritising randomised trials, guidelines and systematic reviews for the clinical questions and validation studies and reviews for the AI applications. The evidence base is uneven: automated fracture detection is the most mature application, whereas AI-supported treatment selection, prediction of loss of reduction, prediction of patient-reported outcomes and rehabilitation triage remain largely investigational in this population. Implementation barriers are substantial, including limited dataset quality, incomplete external validation, difficult workflow integration and the digital exclusion of many older adults. AI is therefore best understood as a potential decision-support layer rather than a replacement for clinical judgement. If rigorously validated, its most realistic near-term contributions are standardising assessment, flagging patients who need closer follow-up and extending follow-up capacity in high-volume environments. By separating established from investigational applications on an explicit readiness gradient, distinguishing prediction of surgical need from prediction of surgical benefit, and separating telerehabilitation from AI-based rehabilitation, this review offers a realistic, implementation-focused framework for using AI in high-volume, resource-constrained fracture care and for prioritising the validation still required.
Sidar Öztürk, Zafer Volkan Gökçe· Bulletin of the National Res...· 0 citations
Background: Artificial intelligence (AI) has expanded rapidly across orthopaedic practice, yet routine clinical adoption remains limited despite strong technical performance. This narrative review examines why a persistent gap separates technical maturity from clinical maturity across the orthopaedic patient care pathway. Methods: We performed a structured qualitative evidence synthesis of contemporary high-level evidence (systematic reviews, diagnostic test accuracy meta-analyses, and structured narrative reviews) retrieved from PubMed/MEDLINE, Scopus, and Web of Science, supplemented by backward screening of reference lists, covering January 2022 to June 2026. Twenty-one evidence syntheses were analysed thematically across six predefined analytical domains and organized according to the orthopaedic patient pathway. Reporting followed the SANRA (Scale for the Assessment of Narrative Review Articles) criteria. Results: Musculoskeletal imaging and fracture detection represented the most mature domains, with several applications reaching early clinical adoption. Applications in arthroplasty planning, shoulder surgery, perioperative prediction, multimodal AI, and clinical decision support remained at developing or emerging stages. Recurrent barriers included limited external validation, dataset heterogeneity, poor workflow interoperability, limited explainability, regulatory and ethical uncertainty, and scarce patient-centred outcome evidence. Conclusions: The principal challenge facing orthopaedic AI is no longer algorithm development but clinical translation. We propose the ORION Clinical Readiness Framework, an evidence-informed five-domain model describing the transition from Technical Performance through Clinical Validation, Workflow Integration, and Patient Benefit to Routine Clinical Adoption, to guide implementation and future research.
Rafael De Nigris González, P. Mello· Journal of Clinical Medicine· 0 citations
Background Artificial intelligence (AI) is increasingly used to enhance diagnostic accuracy, automate image interpretation, and support clinical decision-making. In the field of spine care, applications include MRI and CT-based detection of lumbar disc degeneration, spinal stenosis, vertebral fractures, and axial spondyloarthritis, as well as emerging symptom-based and multimodal diagnostic tools. However, evidence remains dispersed across modalities and conditions, and the quality and clinical readiness of AI systems vary. This scoping review maps current AI applications for diagnosing spinal disorders and identifies gaps for future research and clinical translation. Methods This review followed Joanna Briggs Institute (JBI) and PRISMA-ScR guidelines. Ovid MEDLINE, AMED, Embase, Cochrane CENTRAL, Web of Science, and Scopus were searched from January 2019 to December 2024. Eligible studies were mapped according to AI methodology, diagnostic target, data source, and validation approach, and were required to involve human participants, include sufficient methodological detail, and published in English peer-reviewed journals. No geographic restrictions were applied. Data was extracted on study design, AI methodology, diagnostic target, validation approach, and usability. Methodological quality was assessed using a 19-point scoring system covering study design, reporting clarity, data validation, and feature selection. Results Forty-six studies met the inclusion criteria, conducted primarily in Asia and Europe, with two studies from North America and one from South America. Most investigations were retrospective, imaging-based deep learning models applied to MRI or CT for detecting disc herniation, lumbar spinal stenosis, modic changes, vertebral fractures, and sacroiliitis. Several studies used prospective designs or external validation. Diagnostic performance was generally high across imaging models, with many studies describing accuracy that approached or matched clinician benchmarks, particularly in sacroiliitis classification, disc disease detection, and stenosis grading. Methodological scores ranged from 7.5 to 17.5 out of 19, with recurrent weaknesses in handling missing data, feature selection, and data element validation. Conclusion This review maps a growing body of literature on AI applications for diagnosing spinal disorders, with studies most frequently reporting favorable performance for MRI- and CT-based detection of degenerative and inflammatory conditions. Evidence remains preliminary and heterogeneous.
Victoria A. Bensel, Anne Habeck, Marcda Hilaire Brunot et al.· PLoS ONE· 0 citations
Chronic bone and joint pain is highly prevalent, and rehabilitation is a core component of care. This critical narrative review uses a transparent, structured evidence-mapping approach to examine rehabilitation studies published from January 2021 to April 2026, integrate findings across intervention classes, and identify evidence gaps. Across the included evidence base, exercise therapy supports pain and functional outcomes in selected musculoskeletal conditions, although optimal dose and phenotype-specific matching remain uncertain. Manual therapy and physical modalities may provide short-term adjunctive symptom relief, but their effects are conditional on technique, context, and protocol. Digital rehabilitation and virtual-reality-based approaches can improve access and engagement, but functional and long-term benefits are heterogeneous. Low-load blood-flow-restriction training is promising where high external loads are not tolerated; safety, dose, and longer-term effectiveness require further study. Rather than implying a universal precision-rehabilitation model, we present phenotype-informed treatment matching as a research agenda. Key limitations include heterogeneity of populations, interventions, outcome measures, and follow-up, the selective nature of narrative synthesis, publication bias, and the absence of pooled quantitative estimates. Exercise remains a foundation of care, with adjuncts selected according to presentation, patient priorities, feasibility, and response to reassessment.
Jie Zhuang, Jin-Yan Wang, Yin-Hu Hu et al.· Frontiers in Pain Research· 0 citations