Fusion of CNN and swin transformer for robust lung sound classification
Abstract
Automatic lung sound classification plays a pivotal role in the diagnosis of chronic respiratory diseases, enabling early screening and informed clinical decisions. However, evaluating model performance on imbalanced datasets, such as the ICBHI 2017, remains challenging. As a result, existing studies have focused on improving overall accuracy or average score, while paying limited attention to achieving balanced detection between pathological and healthy events. From a clinical perspective, such an imbalance introduces systematic diagnostic bias, manifested as over-detection or under-detection of disease or abnormal lung sounds, even when high-performance metrics are reported. This study evaluates 1D-CNN, 2D-CNN, Swin Transformer, and a fusion CNN–Swin Transformer model, including its pretrained variant, under the official evaluation protocol. The fusion model combines CNN-based local time-frequency extraction and global contextual modeling using the Swin transformer. The results show that the fusion architecture and its pretrained variant across all ICBHI tasks have a balanced detection between true positive and true negative cases. Improvements in sensitivity up to 5.8% and harmonic score up to 3.5% are observed in certain tasks, reflecting strong generalization in detecting pathological events while maintaining specificity.