A Deep Learning System for Automatic Localization of Anatomical Landmarks in X-rays to Assist in Diagnosis and Surgical Planning
Abstract
Accurate localization of anatomical landmarks is crucial for clinical diagnosis and treatment assessment. However, existing Convolutional Neural Network (CNN)-based methods may result in global spatial information loss and consequent localization failures in the presence of complex anatomical structures or parenchymal abnormalities. Therefore, a method capable of modeling global context while preserving local information is needed. Leveraging the Transformer’s ability to capture long-range dependencies, we propose a novel landmark localization framework, Res-SwinFusion, which integrates a Swin Transformer and a classical CNN backbone in parallel. To effectively merge their complementary features, we designed a feature interactive aggregation module that fuses semantic representations from both branches. Additionally, we introduced a discrimination feature guidance module to provide pixel-level cues and disambiguate landmark locations. We further analyzed the effects of various Gaussian heatmap settings on convergence. Res-SwinFusion achieved strong performance across three anatomical landmark localization datasets. The mean radial errors were 1.04 mm and 1.37 mm on two public cephalogram test sets, 0.63 mm on a public hand X-ray dataset, and 1.44 mm on an internal pelvic X-ray dataset. Ablation studies indicated that Transformer-based global modeling, feature interactive aggregation, and discrimination feature guidance each contributed to improved localization accuracy. The proposed Res-SwinFusion framework offers a solution for anatomical landmark localization with enhanced robustness and precision by combining global contextual modeling and local feature preservation. Code is publicly available at https://github.com/JZK00/Res-SwinFusion .