An Asynchronous Gradient Fusion Framework for Distributed Deep Learning in Tuberculosis Chest X-Ray Diagnosis
Abstract
The rapid increase in high-resolution medical imaging data has highlighted clear limitations in traditional single-node deep learning, especially for time-sensitive healthcare tasks such as Tuberculosis (TB) diagnosis. Although deep learning models perform well in medical image analysis, their high computational requirements often lead to long training times and communication bottlenecks in distributed setups. To address this, we propose a Collaborative Gradient Fusion (CGF), a distributed deep learning framework designed to improve scalability and training stability in asynchronous environments. CGF is built on a Ray-based distributed cluster, where multiple Computational Units train in parallel on separate data partitions, while a Centralized Aggregation Hub coordinates both local and global model updates using a parameter-server design. The framework supports asynchronous gradient fusion through non-blocking, enabling continuous learning without strict synchronization delays. To improve stability, CGF introduces a version-aware aggregation strategy with bounded staleness control, along with adaptive batch-size tuning to reduce the impact of delayed gradients. The aggregation hub uses optimization approaches such as Stochastic Gradient Descent with momentum and ADAM, which support flexible quorum and timeout-based synchronization. Experiments on the TB Chest X-ray dataset and CIFAR-10 using DenseNet121, ResNet101, and MobileNetV2 show that CGF improves accuracy by 8.5–9.5%, boosts F1-score and AUC as compared to the existing Parameter Server Strategy, while also improving Cohen’s Kappa.