Accurate and efficient automated analysis of Optical Coherence Tomography (OCT) images is critical for large-scale retinal disease screening. However, current deep learning models often fail to simultaneously achieve high classification accuracy, practical computational feasibility, and interpretability. To deal with these problems, this paper presents a deep learning framework on YOLOv11 for eight-class retinal disease classification optimized using a two-stage optimization strategy. The proposed architecture incorporates a Progressive Spatial Fusion (PSF) module to hierarchically integrate multi-scale feature representations, followed by a Squeeze-and-Excitation (SE)-based dual-attention refinement mechanism that refines disease-discriminative features prior to classification. Evaluated on the OCT-C8 dataset, the proposed model achieves 98.00% accuracy, precision, recall, and F1-score, together with 99.70% specificity, while achieving perfect classification performance for the Age-related Macular Degeneration (AMD), Central Serous Retinopathy (CSR), Diabetic Retinopathy (DR), and Macular Hole (MH) classes. Grad-CAM visualizations provide qualitative insights into image regions influencing model predictions and show correspondence with disease-associated structural patterns reported in OCT literature. These results show that the proposed framework provides a good trade-off between classification performance, computational requirements and interpretability and shows potential for automated OCT-based retinal disease screening support.
Mithun Vijayan, Goutham Veerapu, N. S.· Frontiers in Artificial Inte...· 0 citations
Accurate classification of retinal diseases from optical coherence tomography (OCT) images is important for clinical decision support. However, many studies in deep learning put the primary focus on enhancing accuracy and often provide limited attention to calibration and robustness under real-world variations. This work presents an enhanced ConvNeXt-Tiny framework incorporating a Spatial Adaptive Multi-scale Convolution (SAMC) module for retinal OCT classification. Rather than introducing a new convolutional operation, SAMC combines established multi-scale dilated convolution and feature fusion principles in a lightweight stage-wise configuration within the ConvNeXt hierarchy to refine representations of retinal abnormalities occurring at different spatial scales. The model is evaluated on the OCT-C8 dataset with eight retinal disease classes. It obtains a test accuracy of 98.39% and a macro F1 score of 0.9839. Reliability of classification was also analyzed by measuring Expected Calibration Error (ECE) and Brier score. Temperature scaling is applied as a
post-hoc
calibration method, while Monte Carlo dropout is used to estimate predictive variability under stochastic inference. Robustness is evaluated under noise, blur, and illumination changes, where minimal performance drop was observed. Computational complexity analysis is performed in terms of model parameters, floating-point operations, and inference efficiency to assess the practicality of the proposed framework. Statistical significance is confirmed using McNemar's test. Interpretability analysis with Grad-CAM and LIME reveals the model attends to clinically significant regions of the retina. Overall, the proposed ConvNeXt-SAMC framework achieves accurate retinal disease classification, while additional calibration, uncertainty, robustness, statistical, and interpretability analyses provide a broader assessment of model behavior.
Mithun Vijayan, Goutham Veerapu, N. S.· Frontiers in Artificial Inte...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.