Skip to content
Conference

Cross-Model Attention-based Vision–Language Fusion for Interpretable Multi-Label Chest Disease Diagnosis

Jul 2026 · 2026 International Conference on Intelligent and Sustainable AI Systems (ICOSAAS) · pp. 1481-1486 · 0 citations · 17 references

Abstract

Chest diseases including pneumonia, cardiomegaly, pleural effusion, atelectasis and consolidation are major public health problems around the world that need to be diagnosed quickly and accurately to provide appropriate treatment. In this paper, we propose a cross-modal attention-based multimodal deep learning method that integrates both chest X-ray (CXR) image and related radiology report for automated multi-label classification of chest diseases. A convolutional network is used as the backbone for the visual representations while semantic textual information is extracted using the BioBERT-based language modeling. To enable meaningful interaction between image and text features and enhance feature integration and diagnostic inference, a cross-modal attention module is introduced. For better explanation of predictions, both visual by Grad-CAM and textual by SHAP are used. Extensive experiments conducted on MIMIC-CXR and NIH ChestX-Ray14 datasets show that the proposed framework outperforms other unimodal and conventional fusion methods by not only accuracy but also AUC, recall, and overall robustness. The results show that multimodal vision–language learning is a clinically applicable, interpretable, and effective approach to computer-aided diagnosis of chest diseases.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.