Aug 2026· Medicina· Vol 62· 0 citations· 45 references
Medicine
TL;DR
A two-stage decomposition that separates systemic skeletal fragility from within-patient vertebral outlier status produces well-calibrated per-vertebra fracture-risk estimates from routine clinical lumbar spine CT and opportunistic abdominal CT.
Abstract
Background and Objectives: Osteoporotic vertebral compression fractures affect approximately one in four postmenopausal women and carry substantial morbidity, yet established clinical tools such as dual-energy X-ray absorptiometry (DXA) provide only patient-level risk and do not identify which specific vertebra is most likely to fail. Computed tomography (CT) acquired for unrelated indications is the most widely available three-dimensional substrate for opportunistic screening, but published machine learning models for vertebral fracture risk almost universally operate at the patient level. The present study aimed to develop and rigorously validate a per-vertebra prediction pipeline applicable to both routine clinical lumbar-spine CT and opportunistic abdominal CT, both acquired for indications unrelated to osteoporosis screening. Materials and Methods: Two independent retrospective cohorts were assembled from a single academic centre: a routine clinical lumbar-spine CT cohort of 106 patients yielding 478 evaluable vertebrae, and a routine abdominal CT cohort of 126 patients yielding 589 evaluable vertebrae. Vertebral bodies were segmented automatically with TotalSegmentator v2 and the trabecular core isolated by morphological erosion. A panel of 505 quantitative imaging biomarkers compliant with Image Biomarker Standardisation Initiative recommendations was extracted, covering trabecular density, vertebral morphometry, classical texture, trabecular network architecture, sub-endplate vulnerability, low-density topology, radial heterogeneity and adjacent muscle quality. Within-patient feature engineering expanded the input pool to 1293 contextual descriptors. Three model families were evaluated under fully nested leave-one-patient-out cross-validation: ElasticNet logistic regression, a softmax-ranking approximation of conditional logistic regression, and a Two-Stage model combining a patient-level fragility score with a within-patient vertebral outlier score. Patient-level bootstrap resampling (2000 iterations) was used to obtain 95% confidence intervals. Results: On routine clinical lumbar-spine CT the Two-Stage model achieved a per-vertebra AUC of 0.750 (95% CI 0.704 to 0.795), an F1 of 0.549, a within-patient concordance index of 0.693, an expected calibration error of 0.044, and Hit@3 of 0.934. It was the only model evaluated that returned calibrated probabilities; the softmax-ranking and ElasticNet baselines gave expected calibration errors of 0.232 and 0.218 respectively. On opportunistic abdominal CT, the softmax-ranking model gave AUC 0.672 (95% CI 0.615 to 0.727). Selected biomarkers were dominated by regional trabecular density and trabecular network architecture; a stable core of lumbar features entered the model in 100% of cross-validation folds, indicating high reproducibility. The closest prior per-vertebra CT-based predictor in primary, non-surgical patients (Muehlematter and colleagues, 58-patient cohort) reported a per-vertebra AUC of 0.64, which is one of several reference points for the present results. Ten methodological variants and sensitivity analyses, including rank fusion, internal tissue normalisation and additional biomechanical features, did not provide statistically significant gains, indicating that the binding constraint at this sample size is data volume rather than methodology. Conclusions: A two-stage decomposition that separates systemic skeletal fragility from within-patient vertebral outlier status produces well-calibrated per-vertebra fracture-risk estimates from routine clinical lumbar spine CT and was the only model evaluated to do so, which is what permits a per-vertebra output to be reported as an absolute risk rather than as an ordering alone; a within-patient ranking model is preferable for opportunistic abdominal CT. The discrimination advantage of the decomposition over that baseline is numerical and consistent but not statistically established at this sample size, and the work is presented as a transparent and reproducible single-centre benchmark for the still under-developed per-vertebra prediction task. Its clearest near-term value is opportunistic, namely flagging elevated per-vertebra fracture risk on CTs already acquired for unrelated indications without additional radiation, cost or a dedicated densitometric study. External multi-centre validation is the necessary next step.
The nomogram developed based on conventional clinical data in this study, has undergone internal validation and demonstrated efficacy in predicting NVCF following PVP, and serves as a valuable decision-making tool for clinicians.
Yi-Qi Wu, Qing Song, Canwei Hu et al.· Frontiers in Medicine· 0 citations
DL analysis of lumbar radiographs, combined with transfer learning and metaheuristic optimization, may support BMD category classification as a complementary screening tool for osteoporosis screening and diagnosis.
Deniz Apalan, Bulent Yildiz, Ozlem Polat et al.· Journal of Bone and Mineral...· 0 citations
Purpose To develop and validate an AI algorithm that enables dual-site (spine and hip) BMD assessment for opportunistic osteoporosis screening using a single kidney-ureter-bladder (KUB) radiograph. Materials and methods In this institutional review board approved prospective study, we developed the SHield pipeline for opportunistic osteoporosis screening. From an initial dataset of 15,175 KUB images, a final cohort of 4,436 patients (mean age 69.0 ± 12.2 years) was included for model development after exclusions for suboptimal quality or incomplete region of interests. The pipeline analyzes both the spine and hip regions on a single KUB to predict site-specific BMD values. Finally, it integrates these dual-site predictions to determine the lowest T-score, mimicking the standard clinical DXA reporting protocol which evaluates both the lumbar spine and proximal femur. Results On the internal test set (628 patients), the hip and spine AI models demonstrated Pearson correlation coefficients of 0.887 and 0.921 and RMSEs of 0.057 and 0.062, respectively, compared to DXA-measured BMD, achieving AUCs of 0.894 and 0.956 for predicting T-scores ≤-2.5. In a subsequent prospective feasibility study (51 patients), we employed a conservative detection threshold of Tm-score ≤-2.8 to minimize false positives and unnecessary referrals. The final AI pipeline achieved a Pearson correlation of 0.959, an AUC of 0.969, 100% PPV, 100% specificity, and 83.8% NPV. Conclusion This paper presents the first AI pipeline that enables gold-standard-aligned, dual-site BMD estimation and opportunistic osteoporosis screening from a single KUB radiograph. Validated prospectively, its high accuracy and specificity confirm its potential to significantly increase early detection rates without overwhelming clinical workflows.
Hsuan-Yin Lin, J. Chai, Qingzong Tseng et al.· Frontiers in Endocrinology· 0 citations
OBJECTIVE
To develop and validate a hierarchical deep learning model for differentiating acute and chronic lumbar osteoporotic vertebral compression fractures (OVCFs) using X-ray images.
MATERIALS AND METHODS
We retrospectively reviewed approximately 2600 lateral lumbar radiographs obtained from patients clinically suspected of having OVCFs between 2007 and 2022. After excluding poor-quality images and surgically instrumented vertebrae, 1299 radiographs (6495 vertebral patches, L1-L5) were included. Labeling was performed by neurosurgeons and radiologists using X-ray images, with CT and/or MRI findings serving as the reference standard. A two-step hierarchical classification was implemented: first classifying vertebrae into Normal-Chronic, Acute, and Indeterminate (cement-augmented vertebrae without instrumentation) groups, followed by subdivision of the Normal-Chronic group into Normal and Chronic categories.
RESULTS
A total of 1299 radiographs were evaluated. The hierarchical model achieved an accuracy of 91% in the initial three-class step. For the detection of acute fractures in the final classification step, the model demonstrated a sensitivity of 91.0% (95% CI 84.8-95.0%), a specificity of 82.1% (95% CI 79.8-84.5%), and a high negative predictive value (NPV) of 98.8% (95% CI 97.9-99.3%). The Normal-aligned hierarchical approach outperformed the Acute-aligned and end-to-end models, particularly for acute and chronic cases.
CONCLUSION
The proposed hierarchical approach enhances the diagnostic utility of standard X-ray images by enabling more accurate classification of lumbar fracture types. This study is limited by its single-institution retrospective design. This model may reduce reliance on advanced imaging and support faster and more informed clinical decision-making.
Joohyun Kim, Keewon Shin, Sungjae An et al.· Skeletal Radiology· 0 citations
Background Osteoporotic vertebral compression fractures (OVCFs) are a common spinal disease. Differentiating acute and chronic fractures is the key to determining the treatment plan. To develop and evaluate a deep learning model capable of differentiating acute and chronic OVCFs on CT images. Methods The internal dataset comprised CT images of 624 fractured vertebrae from 400 patients with OVCF treated at Hospital 1 between January 1, 2020, and November 1, 2023. The patients were randomly divided into a training set (340 patients, 532 fractured vertebrae, 85.0%) and an internal testing set (60 patients, 92 fractured vertebrae, 15.0%). An external testing set included CT images of 86 fractured vertebrae from 70 patients with OVCF treated at Hospital 2 from January 1, 2023, to November 1, 2023, using identical inclusion criteria. Three radiologists manually delineated regions of interest (ROIs) in CT images of diseased vertebrae using Anaconda Prompt software with rectangular bounding boxes. The trained YOLO v7 model’s performance was evaluated on both the internal and external testing sets and compared with that of the three radiologists using metrics with accuracy, sensitivity, specificity, and AUC. Results The internal testing set revealed the following: AUC, 0.937 (95%CI: 0.932–0.991); accuracy, 0.961 (0.930–0.977); sensitivity, 0.973 (95%CI: 0.958–0.994); specificity, 0.761 (95%CI: 0.728–0.784); precision, 0.936 (95%CI: 0.920–0.966), and F1 score; 0.954 (95%CI: 0.913–0.975). The external testing set revealed the following: AUC, 0.882 (95%CI: 0.839–0.894); accuracy, 0.852 (95%CI: 0.827–0.886); sensitivity, 0.895 (95%CI: 0.840–0.913); specificity, 0.780 (95%CI: 0.730–0.812); precision, 0.874 (95%CI: 0.839–0.892); and F1 score, 0.884 (95%CI: 0.846–0.903). Conclusion The YOLO v7 model achieved good performance for CT-based differentiation of acute versus chronic OVCFs in patients and showed better performance compared to radiologists in both testing sets, which may serve as a decision-support tool for CT-equivocal cases.
AI-assisted proximal femur BMD assessment on contrast-enhanced abdominal CT showed high agreement and classification concordance with unenhanced CT-derived quantitative computed tomography (QCT), although external validation against DXA or independent QCT remains necessary.
Fenghuang Lin, Juan Xu, Yan-Xia Chen et al.· Frontiers in Medicine· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.