Skip to content
Open access

A two-stage deep learning framework with progressive feature fusion and attention mechanisms for multi-class retinal disease classification

Sep 2026 · Frontiers in Artificial Intelligence · 0 citations · 57 references

Abstract

Accurate and efficient automated analysis of Optical Coherence Tomography (OCT) images is critical for large-scale retinal disease screening. However, current deep learning models often fail to simultaneously achieve high classification accuracy, practical computational feasibility, and interpretability. To deal with these problems, this paper presents a deep learning framework on YOLOv11 for eight-class retinal disease classification optimized using a two-stage optimization strategy. The proposed architecture incorporates a Progressive Spatial Fusion (PSF) module to hierarchically integrate multi-scale feature representations, followed by a Squeeze-and-Excitation (SE)-based dual-attention refinement mechanism that refines disease-discriminative features prior to classification. Evaluated on the OCT-C8 dataset, the proposed model achieves 98.00% accuracy, precision, recall, and F1-score, together with 99.70% specificity, while achieving perfect classification performance for the Age-related Macular Degeneration (AMD), Central Serous Retinopathy (CSR), Diabetic Retinopathy (DR), and Macular Hole (MH) classes. Grad-CAM visualizations provide qualitative insights into image regions influencing model predictions and show correspondence with disease-associated structural patterns reported in OCT literature. These results show that the proposed framework provides a good trade-off between classification performance, computational requirements and interpretability and shows potential for automated OCT-based retinal disease screening support.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.