Skip to content
Review Open access

A trustworthy cross-domain AI framework for fundus disease classification using hybrid CNN fusion and supervised domain adaptation

Aug 2026 · Discover Artificial Intelligence · Vol 6 · 0 citations · 45 references

TL;DR

FusionEye-Net combines strong internal performance with substantially improved cross-domain generalization through lightweight adaptation, underscoring the critical role of external validation and domain-aware fine-tuning in formulating trustworthy, adaptive AI-assisted decision-support tools that can safely augment automated ophthalmic screening workflows.

Abstract

Deep learning–based systems for fundus disease classification often achieve impressive accuracy on internal datasets, yet their performance degrades markedly when applied to real-world clinical data. This limitation is primarily caused by domain shift arising from variations in imaging devices, illumination conditions, and population characteristics, which remains a key barrier to integrating these technologies into reliable clinical decision-support systems. To address this challenge, we propose FusionEye-Net, a hybrid deep learning framework that performs feature-level fusion of EfficientNet-B3 and ResNet-50 for four-class fundus image classification, namely cataract, diabetic retinopathy, glaucoma, and normal retina. A dedicated preprocessing pipeline—including circular fundus cropping, illumination normalization, contrast-limited adaptive histogram equalization (CLAHE), and adaptive gamma correction—was employed to reduce inter-device variability and standardize retinal appearance. FusionEye-Net was first trained on a curated internal dataset of approximately 4000 images, achieving an internal test accuracy of 99.24%. However, evaluation on an independent external dataset comprising 4640 images revealed a significant performance drop to 72.05%, highlighting the severity of cross-domain variability. To mitigate this degradation, a targeted supervised domain adaptation strategy was applied using a balanced subset of 2000 external images (500 per class). This adaptation improved external performance to 91.27% accuracy with a macro F1-score of 0.9139, reducing misclassifications by more than two-thirds. Model predictions and visual explanations were reviewed by a board-certified retina specialist to ensure clinical plausibility. Explainability analysis using Gradient-weighted Class Activation Mapping (Grad-CAM) and region-of-interest (ROI) contour mapping demonstrated that the adapted model consistently focuses on medically relevant structures, such as lens opacity in cataract, microaneurysms in diabetic retinopathy, and optic disc cupping in glaucoma. In summary, the results demonstrate that FusionEye-Net combines strong internal performance with substantially improved cross-domain generalization through lightweight adaptation, underscoring the critical role of external validation and domain-aware fine-tuning in formulating trustworthy, adaptive AI-assisted decision-support tools that can safely augment automated ophthalmic screening workflows.

Read PDF

Similar papers

Open access Sep 2026

Improving Generalization of Deep Learning for Glaucoma Classification Under Real-World Domain Shift via Unsupervised Domain Adaptation

Purpose To develop and evaluate an unsupervised domain adaptation (UDA) framework for glaucoma classification from fundus images that improves the generalizability of deep learning (DL) models across heterogeneous imaging characteristics and clinical settings. Methods We developed an adversarial UDA framework that adapts a labeled source domain to an unlabeled target domain by jointly optimizing glaucoma classification and domain discrimination. A total of 6906 fundus images were included: 1422 images (full-view and optic nerve head-cropped) derived from 711 fundus photographs of 520 patients from the University of Illinois Chicago and 5484 images from three public datasets (RIMONE-DL, n = 371; REFUGE, n = 259; LAG, n = 4854). Performance was evaluated across multiple source–target domain pairs. Saliency map analysis compared feature utilization between UDA and standard DL models. Results Across diverse source–target dataset pairs, the proposed UDA framework improved standard DL classification accuracy by up to 40.0% and area under the curve (AUC) by up to 59.8% in extreme domain shift experiments. UDA also reduced the performance gap to the ideal target-trained model by up to 82.6% for accuracy and 38.1% for AUC. When applied to unseen domains, UDA improved classification accuracy by up to 14.7% relative to standard DL. Conclusions The proposed UDA framework improves cross-domain generalizability of DL-based glaucoma classification from fundus images. Compared with standard DL models, UDA exhibits feature utilization patterns more aligned with an ideal baseline, reducing overreliance on optic nerve head–specific cues and supporting more robust decision-making under domain shift. Translational Relevance By enabling glaucoma classification without requiring labeled data from new clinical sites, this UDA framework addresses a key barrier to deploying artificial intelligence–based screening tools across diverse real-world ophthalmic settings.

Homa Rashidisabet, R. V. Paul Chan, T. Vajaranant et al. · 0 citations
Jul 2026

High Performance Hybrid CNN CBAM Framework for High Sensitivity Heart Disease Classification

Cardiovascular diseases remain a paramount global health crisis, necessitating early and precise diagnostic interventions. While medical imaging is the clinical standard, manual interpretation is highly susceptible to visual fatigue and inter-observer variability. This study proposes a novel, highly robust Computer-Aided Diagnosis (CAD) framework that overcomes the spatial and textural limitations of standalone Convolutional Neural Networks (CNN) in heart disease image classification. A Feature-Level Ensemble (Hybrid) architecture was created by putting together the deep semantic features of ResNet50V2, the spatial boundaries of VGG16, and the parameter efficiency of EfficientNetV2B3. To directly deal with the loss of features caused by anatomical background noise, a Convolutional Block Attention Module (CBAM) was added to the EfficientNet pathway. This gave the network two-dimensional (channel and spatial) visual attention. To guarantee a thorough and impartial assessment, a complete restructuring of a dataset comprising 5,977 images was undertaken using an 80:10:10 stratified split, thereby eliminating the accuracy paradox resulting from class imbalance. The proposed Hybrid CBAM model significantly outperforms standalone baselines, with a peak accuracy of 94.00%. For clinical use, it was very important that the attention-guided ensemble had a Recall (sensitivity) of 0.94 for finding pathological cases and a Negative Precision of 0.96. This study definitively demonstrates that the integration of multi-model feature extraction with focused visual attention mechanisms yields a highly sensitive, reliable, and non-invasive automated screening instrument for the early detection of cardiovascular disease.

Giant Prakoso Amukti Wibowo, Slamet Riyadi, A. Dewi et al. · 0 citations
Open access Sep 2026

PIKER-NET: Multi-class retinal disease classification using Pied Kingfisher optimization-based improved residual network

Retinal diseases are vision-threatening conditions, including age-related macular degeneration (ARMD), diabetic retinopathy (DR), and glaucoma, that require early and accurate detection to prevent blindness. However, existing methods often struggle with limited feature representation, high inter-class similarity, intra-class variability, and reduced performance in handling noisy and low-quality retinal images. To address these challenges, a novel PIKER-NET framework is proposed for accurate multi-class retinal disease classification. The input fundus images from the RFMiD dataset are pre-processed using a scalable range adaptive bilateral (SCRAB) filter to enhance image clarity by preserving edges while reducing noise. The Improved Residual Network-Rescaled (ImResNet-RS) integrated with Temporal Attention is then employed to extract deep hierarchical features with enhanced discriminative power. Pied Kingfisher Optimization (PKO) algorithm is utilized for feature selection, effectively reducing redundant information while retaining the most relevant features. Residual Multilayer Perceptron (ResMLP) is used to classify retinal diseases into ARMD, branch retinal vein occlusion (BRVO), diabetic neuropathy (DN), DR, healthy, and myopia (MYA). The PIKER-NET achieves an overall accuracy of 98.14% and F1-score of 97.06%. The PIKER-NET approach improves overall accuracy by 3.24%, 4.24%, 6.23%, and 2.00% compared to EyeDeep-Net, IDL-MRDD, DeepDiabetic, and VisionDeep-AI, respectively. The proposed approach has strong clinical relevance by supporting earlier disease screening, reducing misdiagnosis, and enabling faster diagnosis to assist ophthalmologists in improving patient outcomes.

Lissy Devasahayam, Ramya Devi Murugadasan, Anandhi Samuel Vijayalakshmi et al. · 0 citations
Open access Sep 2026

A two-stage deep learning framework with progressive feature fusion and attention mechanisms for multi-class retinal disease classification

Accurate and efficient automated analysis of Optical Coherence Tomography (OCT) images is critical for large-scale retinal disease screening. However, current deep learning models often fail to simultaneously achieve high classification accuracy, practical computational feasibility, and interpretability. To deal with these problems, this paper presents a deep learning framework on YOLOv11 for eight-class retinal disease classification optimized using a two-stage optimization strategy. The proposed architecture incorporates a Progressive Spatial Fusion (PSF) module to hierarchically integrate multi-scale feature representations, followed by a Squeeze-and-Excitation (SE)-based dual-attention refinement mechanism that refines disease-discriminative features prior to classification. Evaluated on the OCT-C8 dataset, the proposed model achieves 98.00% accuracy, precision, recall, and F1-score, together with 99.70% specificity, while achieving perfect classification performance for the Age-related Macular Degeneration (AMD), Central Serous Retinopathy (CSR), Diabetic Retinopathy (DR), and Macular Hole (MH) classes. Grad-CAM visualizations provide qualitative insights into image regions influencing model predictions and show correspondence with disease-associated structural patterns reported in OCT literature. These results show that the proposed framework provides a good trade-off between classification performance, computational requirements and interpretability and shows potential for automated OCT-based retinal disease screening support.

Mithun Vijayan, Goutham Veerapu, N. S. · 0 citations
Open access Sep 2026

Comparative Analysis of CNN and Transformer Models for Multi-Class Diabetic Retinopathy Grading Using Fundus Images

Background/Objectives: Diabetic retinopathy is a major cause of preventable vision loss worldwide, making early and accurate disease grading crucial for timely treatment. Although convolutional neural network (CNN)- and transformer-based architectures have demonstrated promising performance for retinal image analysis, comprehensive comparisons within a unified experimental framework remain limited. This study systematically compares representative standard and lightweight CNN- and transformer-based architectures for multi-class DR grading. Methods: Six ImageNet-pretrained deep-learning models, including ResNet50, EfficientNet-B0, MobileNetV2, Vision Transformer (ViT), Swin-Tiny, and Swin Transformer, were evaluated on the APTOS 2019 retinal fundus image dataset under a unified experimental configuration with consistent preprocessing, data augmentation, training, and evaluation settings. All models were fine-tuned and evaluated independently over five runs with different random seeds. Their performance was assessed using accuracy, precision, recall, F1-score, area under the receiver operating characteristic curve (AUC), Quadratic Weighted Kappa (QWK), per-class analysis, computational efficiency, and statistical analysis. Results: Transformer-based models generally achieved higher mean classification performance than the evaluated CNN-based models. Swin-Tiny achieved the highest mean accuracy (82.3%), macro F1-score (64.4%), weighted F1-score (82.1%), and QWK (89.8%) across the five runs. Among the CNN-based models, EfficientNet-B0 achieved the strongest overall classification performance, whereas MobileNetV2 provided the lowest computational complexity. The results also highlighted differences in learning behavior and computational requirements across the evaluated architectures. Repeated experiments demonstrated stable performance across different random seeds, supporting the reliability of the proposed evaluation. Conclusions: Overall, this study provides a comprehensive comparison of representative CNN- and transformer-based architectures under consistent experimental settings and offers practical guidance for selecting suitable deep learning models for automated diabetic retinopathy screening.

Maha A. Thafar · 0 citations
Review Open access Aug 2026

AI-Based Skin Disease Classification Using Deep Learning Techniques

Skin diseases represent one of the most widespread categories of health disorders worldwide, and timely diagnosis plays a critical role in preventing complications such as skin cancer. Conventional diagnostic procedures depend largely on visual examination by dermatologists, a process that is subjective, time-consuming, and difficult to access in rural or under-resourced regions. This paper presents a comprehensive deep learning-based framework for the automated detection and classification of skin diseases from dermoscopic and clinical images. The proposed system employs a transfer-learning approach built on the ResNet50 convolutional neural network architecture, pre-trained on ImageNet and fine-tuned on benchmark dermatological datasets, namely HAM10000, the ISIC Archive, and DermNet. The methodology encompasses dataset collection, image pre-processing, data augmentation, feature extraction, model training, disease classification, and rigorous performance evaluation. ResNet50 is selected for its residual-learning capability, which mitigates the vanishing-gradient problem and enables deeper, more accurate networks suited to fine-grained medical image analysis. An extensive review of over thirty related studies spanning convolutional architectures, ensemble methods, and emerging transformer-based models is used to position the proposed framework within the current state of the art. The framework is evaluated using standard classification metrics, including accuracy, precision, recall, specificity, F1-score, Cohen’s kappa, and confusion-matrix analysis, and is benchmarked conceptually against alternative architectures such as a baseline CNN, VGG16, MobileNetV2, DenseNet, and EfficientNet. The anticipated outcome is an accurate, scalable, and accessible screening tool capable of assisting healthcare professionals in early diagnosis, thereby reducing diagnostic delay and improving healthcare accessibility, particularly in regions with limited dermatological expertise.

Nisha Rajodiya, Shailendra Mishra, Sumitra Menaria et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.