Skip to content
Conference

A Hybrid Deep Learning Architecture Integrating YOLOv8 and ResNet for Lung Disease Diagnosis

Jul 2026 · 2026 6th International Conference on Inventive Computation and Information Technologies (ICICIT) · pp. 936-942 · 0 citations · 12 references

Abstract

This project introduces an AI-powered diagnostic platform designed to identify major respiratory conditions from chest X-ray images using a advanced hybrid Deep Learning architecture. By integrating YOLOv8 for precise lesion localization and ResNet50 for deep feature extraction, the system overcomes the limitations of traditional single-model approaches, offering a more detailed analysis of lung pathology. The model is trained on a comprehensive dataset encompassing four critical categories: Normal, COVID-19, Pneumonia, and Tuberculosis. To ensure clinical reliability, the system employs advanced preprocessing including normalization and augmentation to handle variations in X-ray quality. This dual-network engine is integrated into a responsive web application that provides healthcare providers with near-instantaneous diagnostic results and confidence scores. With a user-friendly interface designed for both specialists and general practitioners, the platform bridges the gap in medical expertise, particularly in resource-limited or remote regions. By combining automated detection with accessible web technology, this research provides a scalable solution to accelerate clinical decision-making and improve patient outcomes in respiratory healthcare.

View source

Similar papers

Open access Jul 2026

XRAI: A Deep Learning Approach to Chest Disease Detection

The research introduces a dimension in which medical specialists can have real-time conversations with a trained LLM model that leverages knowledge from a medical encyclopedia, and enhances collaboration between AI and medical professionals, creating a platform for exchanging knowledge.

M. I. Ahmed · 1 citation
Open access Aug 2026

XAI-Enhanced Hybrid CNN–Transformer Framework For Multi-Class Lung Disease Classification

Chest radiograph images have become a critical research area for applying deep learning in radiological interpretation for the classification of pulmonary diseases. But, to achieve both high accuracy and good interpretability continues to be a major hurdle for many researchers. In this research, we offer a hybrid architecture that incorporates CNNs and Transformer techniques for classifying different respiratory diseases using chest radiograph images. The CNN component provides a mechanism to capture many of the fine, local details found in an image, while the Transformer provides a self-attentive mechanism to capture the overall context of an X-ray image. In addition, a range of approaches exist to improve overall performance of the CNN and Transformer architecture, including structured preprocessing, data augmentation and class balancing. All of these techniques will improve model learning performance and help to effectively manage class imbalance when dealing with imbalanced datasets. To make our model more transparent to users and clinically useful, we employed explainability methods like Grad-CAM and Attention Visualizations to provide users with evidence of the specific area in an X-ray where the model is basing its prediction, thereby providing a greater amount of trust on the part of radiologists in interpreting the model's output. Based on our findings from testing the 6 Classes Chest Xray dataset, the proposed system proved to achieve a very impressive final testing accuracy of 94.42%. It classifies tuberculosis and healthy patients particularly well, with precision, recall, and F1-scores of 0.99 and 0.97, respectively, but provides good performance across the other disease types too. Furthermore, confidence analysis of predicted labels exhibited that when there was an accurate prediction, the assigned probability score was usually much higher than the assigned probability score for an incorrect prediction. Thus, these results suggest that the hybrid CNN-Transformer model provides a strong level of diagnostic accuracy and meaningfully understood visual rationale so it can serve as an excellent decision support mechanism for hospitals and radiologists in their daily operations.

Prasanna Pabba, N. S. Chaitanya, M. Ravikanth et al. · 0 citations
Conference Open access 2026

A framework focused on deployment for pulmonary disease classification multimodal deep learning

Globally, pulmonary diseases are a major health burden, especially in settings where resources are constrained and access to specialized radiological expertise is limited. Most chest radiograph-based deep learning models rely only on imaging data, despite showing promising performance when it comes to their diagnostic capabilities. However, clinical decision-making in the real world integrates structured patient information with the aforementioned imaging data. Our study aims to compare current multimodal machine learning approaches that combine clinical data and imaging used in pulmonary disease classification and propose a deployable framework designed for clinical settings in resource-constrained environments. We conducted a review of recent literature around multimodal AI in pulmonary diseases, and we focused on fusion strategies (early-stage, late-stage, and hybrid), techniques for data integration, validation settings, as well as deployment considerations. We performed a comparative synthesis aiming to identify methodological patterns, translational gaps, and performance trends, and on this basis, as a conceptual blueprint to guide our subsequent model development, we propose a modular architecture integrating structured-data encoders with convolutional neural networks for imaging data, along with fusion mechanisms to handle incomplete modalities. We then formalize the fusion operators and present an algorithm for graceful degradation under missing clinical data; empirical validation of the proposed architecture is deferred to a future work. This review indicates that multimodal approaches consistently outperform unimodal imaging models, especially in early-stage or complex cases, with gains in performance reported across many pulmonary conditions. However, the review also reveals several limitations, whether it's the lack of standardized fusion evaluation, insufficiencies in external validation, inadequacies in the handling of missing clinical variables, or limited attention to real-world clinical integration. These obstacles expose the need for system design that is deployment-aware more so than solely performance-driven optimization. Through this work, we aim to contribute a structured synthesis of multimodal pulmonary AI and outline an interpretable, resource-conscious framework intended for integration into healthcare workflows, which we put forward as the design basis for a model to be developed and evaluated in future work. By aligning the model design philosophy with the realities of clinical workflows, this approach should support equitable access to AI-assisted diagnostics and help advance the application of AI for healthcare improvement and social good.

Hamza Hrid, M. Machkour, Y. Asimi · 0 citations
Conference Aug 2026

An Explainable CBAM Enhanced DenseNet121 Framework for Multi-Class Lung Cancer Classification Using CT Scans

Due to its late identification and challenging diagnosis, lung cancer continues to be one of the top causes of death for cancer patients globally, positioning it as one of the most critical concerns. Timely identification of cancerous nodules is essential for enhancing the patient’s survival likelihood CT image analysis by hand is not very productive and significantly relies on a specialist’s expertise. In this study, we offer an autonomous lung cancer classification method based on explainable deep learning. The popular DenseNet121 network serves as the foundation for our deep learning model, which is enhanced by the Convolutional Block Attention Module (CBAM). To improve feature extraction of significant spatial and channel properties of input data, attention techniques are added. Furthermore, our method is interpretable because the Grad-CAM technique makes it possible to explain the choices made by a machine learning system. A database of CT scans, comprising 4,598 pictures categorized by large cell carcinoma, adenocarcinoma, and healthy lungs, was utilized. Our evaluations show the model’s effectiveness with an accuracy rate of 94.6\%.

S. Jegadeesan, S. Matheswaran, R. Palanivelrajan · 0 citations
Open access Aug 2026

An Integrated Computational Framework for Lung Cancer Detection: Combining Deep Learning, Interactive Visualization, and Diagnostic Reporting from CT Imaging

Lung cancer remains the leading cause of cancer-related mortality worldwide, necessitating early and accurate diagnostic solutions. This paper presents an integrated computational framework for lung cancer detection from CT imaging, combining deep learning models with interactive visualization and automated diagnostic reporting. The proposed system leverages two state-of-the-art architectures—EfficientNet with CBAM attention enhancement and Vision Transformer including Swin Transformer variants—to achieve robust classification performance. EfficientNet provides parameter-efficient feature extraction through its compound scaling strategy, while Vision Transformers capture global contextual relationships via self-attention mechanisms, addressing the inherent limitations of CNNs in modeling long-range dependencies in medical images. The framework integrates a hybrid feature fusion approach, interactive visualization modules for clinician interpretability, and automated diagnostic report generation. Experimental evaluation on benchmark datasets demonstrates superior performance, with the optimized CBAM-EfficientNet achieving 99.81% accuracy and the ViT-based fusion approach achieving 99.28% accuracy. The system's interactive visualization capabilities, including Grad-CAM attention maps, enhance clinical interpretability and trustworthiness. This research contributes a comprehensive end-to-end solution bridging computational innovation with clinical practice for improved lung cancer diagnosis.

S. Thilagavathi, R. Ravindran · 0 citations
Conference Aug 2026

A deep learning framework for COVID-19 detection: integrating attention modules into ResNet50 and optimizing cross-entropy loss function

Rapid and reliable diagnosis of COVID-19 is of fundamental importance for effective pandemic control, and chest X-ray imaging combined with deep learning provides a very useful tool for large-scale screening. Therefore, this paper presents a tri-class classification framework for distinguishing COVID-19, viral pneumonia, and normal cases from chest X-rays. The method is built on a pre-trained ResNet50 and incorporates CBAM attention modules at two well-justified stages: spatial attention in Layer3 for abnormal region localization and channel attention in Layer4 for semantic feature selection. More importantly, it addresses the class imbalance common in medical data by using a weighted cross-entropy loss function. Experiments on a public COVID-19 chest X-ray dataset demonstrate that the full model attains 90.74% precision and 74.65% F1-score, both superior to baseline methods. Ablation studies rigorously validate each component, and Precision-Recall curve analysis gives an AUC-PR of 0.942. From the results for the COVID-19 class it is clearly seen that the attention map visualization shows where the model is looking at clinically relevant lung regions, and since the inference time is 2.64 ms per image (379.1 FPS), the proposed framework thus achieves a good balance between accuracy and speed for clinical use.

Weizhen Yu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.