This work proposes a novel framework, RetiWave-Mamba, which integrates spatial-frequency domain learning with state-of-the-art state space models and achieves a state-of-the-art classification accuracy of 98.25%, surpassing existing methods.
Abstract
Retinal diseases are a leading cause of irreversible vision impairment, making early and accurate diagnosis essential for effective treatment. Optical Coherence Tomography (OCT) serves as a critical imaging modality for this purpose, yet its automated analysis is hindered by inherent speckle noise, varying lesion scales, and subtle inter-class similarities. To address these challenges, we propose a novel framework, RetiWave-Mamba, which integrates spatial-frequency domain learning with state-of-the-art state space models. The framework utilizes Discrete Wavelet Transform (DWT) to decompose OCT images into low- and high-frequency streams, enabling decoupled processing of structural context and fine-grained details. For the low-frequency branch, we design a Multi-scale Contextual Localization Module (MCLM), which synergizes multi-scale dilation with spatial attention to expand the global receptive field and precisely localize lesion regions. For the high-frequency branch, we introduce an Attention-Guided High-Resolution Network (AG-HRNet) equipped with an intelligent gating mechanism to suppress noise propagation during multi-scale interactions. Furthermore, a Frequency-Adaptive Mamba Projector (FAMP) is incorporated to capture long-range dependencies within disjoint high-frequency textural features. Extensive experiments on the OCT-C8 dataset demonstrate that our approach achieves a state-of-the-art (SOTA) classification accuracy of 98.25%, surpassing existing methods. These results highlight the efficacy of RetiWave-Mamba in robustly identifying retinal pathologies under noisy conditions, offering a promising tool for clinical diagnosis.
Background Retinal diseases are a major cause of preventable visual impairment, and optical coherence tomography (OCT) provides high-resolution cross-sectional imaging for retinal assessment. However, reliable interpretation of OCT B-scans remains challenging because of speckle noise, subtle layer-wise changes, and visually similar pathological patterns. Methods We propose SLight-Net, a lightweight spectral-layer aware network for retinal OCT classification. The model is built on a compact three-stage convolutional backbone and incorporates two OCT-specific modules. The Frequency-Aware Spectral-Spatial Encoder (FASE) integrates local convolution, dilated contextual modeling, and learnable spectral modulation to capture multi-scale structural and frequency-aware retinal features. The Retinal Layer Depth Attention (RLDA) module further introduces a depth-direction anatomical prior to recalibrate feature responses along retinal layers. Deep supervision and exponential moving-average weight updating are used to improve optimization stability. Results On the OCT-C8 benchmark, SLight-Net achieves 98.21% classification accuracy with only 1.204M parameters. Additional evaluation on OCT2017 shows 99.30% accuracy, suggesting that the model maintains stable performance under a different class setting while remaining compact. Conclusion These findings indicate that OCT-specific spectral and layer-aware priors can support efficient retinal disease classification without relying on large generic backbones, providing a practical basis for lightweight computer-aided OCT analysis.
Yixiang Yao, Jin Hong, Rongli Zhang et al.· Frontiers in Cell and Develo...· 0 citations
To address the significant challenges in brain tumor magnetic resonance imaging detection—including high morphological heterogeneity of lesions, blurred boundaries, and severe background noise interference—this study proposes a multi-domain cooperative perception model, BTA-DETR (Brain Tumor Aware-DEtection TRansformer). First, we design a content-aware fusion unit, which leverages parallel spatial and channel branches alongside an adaptive gating mechanism to dynamically re-weight local textural details and global semantic features, thereby adaptively capturing the diverse morphologies of lesions. Second, we propose the learnable temperature attention module, which generates a pixel-level spatially heterogeneous temperature field via a lightweight convolutional branch to differentially modulate the sharpness of the attention distribution, thereby improving the model’s localization performance in lesion regions with blurred boundaries. Finally, we construct a frequency-spatial cooperative module, which parses pathological textures through a tri-band gating mechanism in the frequency-domain branch, complemented by asymmetric depth-wise directional convolutions in the spatial-domain branch to supplement geometric boundary information, enhancing the representational capacity for fine-grained lesion features. Experimental results on the public Roboflow Brain Tumor Detection dataset demonstrate that, compared to the baseline RT-DETR, BTA-DETR improves mean average precision50 by 4.36 percentage points while reducing parameters by 0.82 M and computational cost by approximately 9.1%. Experimental results on the two external datasets, Kaggle BrainTumor and BraTS 2021 T1ce, further demonstrate that, under a dataset-specific retraining setting, BTA-DETR achieves higher detection metrics than the vanilla RT-DETR baseline.
Liu Fan, Jincheng Zhao, Yuzi Jin et al.· Biomedical engineering and p...· 0 citations
Accurate classification of retinal diseases from optical coherence tomography (OCT) images is important for clinical decision support. However, many studies in deep learning put the primary focus on enhancing accuracy and often provide limited attention to calibration and robustness under real-world variations. This work presents an enhanced ConvNeXt-Tiny framework incorporating a Spatial Adaptive Multi-scale Convolution (SAMC) module for retinal OCT classification. Rather than introducing a new convolutional operation, SAMC combines established multi-scale dilated convolution and feature fusion principles in a lightweight stage-wise configuration within the ConvNeXt hierarchy to refine representations of retinal abnormalities occurring at different spatial scales. The model is evaluated on the OCT-C8 dataset with eight retinal disease classes. It obtains a test accuracy of 98.39% and a macro F1 score of 0.9839. Reliability of classification was also analyzed by measuring Expected Calibration Error (ECE) and Brier score. Temperature scaling is applied as a
post-hoc
calibration method, while Monte Carlo dropout is used to estimate predictive variability under stochastic inference. Robustness is evaluated under noise, blur, and illumination changes, where minimal performance drop was observed. Computational complexity analysis is performed in terms of model parameters, floating-point operations, and inference efficiency to assess the practicality of the proposed framework. Statistical significance is confirmed using McNemar's test. Interpretability analysis with Grad-CAM and LIME reveals the model attends to clinically significant regions of the retina. Overall, the proposed ConvNeXt-SAMC framework achieves accurate retinal disease classification, while additional calibration, uncertainty, robustness, statistical, and interpretability analyses provide a broader assessment of model behavior.
Mithun Vijayan, Goutham Veerapu, N. S.· Frontiers in Artificial Inte...· 0 citations
DSF-Net provides a robust framework for improving vessel continuity and boundary delineation in fundus images and produces more accurate and structurally coherent segmentation results, especially for thin and complex vessels.
Feng Liang, Xiaoqi Sheng, Yang Liu et al.· Frontiers of Computer Scienc...· 0 citations
Retinal vessel analysis in ultra-widefield (UWF) images provides a unique opportunity for large-scale assessment of systemic microvascular health. However, accurate segmentation in true-color UWF images remains challenging due to the large field of view, complex background, and reduced vessel contrast. To address these challenges, we develop ECS-Net, a dedicated deep learning (DL) framework for retinal vessel segmentation in true-color UWF images. Our ECS-Net adopts an enhanced encoder decoder architecture that integrates a dual-domain context enhancement module (DCEM), consisting of an anisotropic spatial focus unit (ASFU) and a feature correlation calibration unit (FCCU), together with atrous spatial pyramid pooling (ASPP) for multi-scale context modeling. The proposed ECS Net achieved a higher Dice coefficient of 0.8349 than other state-of-the-art algorithms (0.7001-0.8136), demonstrating strong generalization under real-world imaging conditions. Building upon accurate vessel extraction, region-specific vascular parameters, including vessel density (VD), fractal dimension (FD), tortuosity (TC), and mean curvature (MC), were quantified separately for central (45°), peripheral (45° 133°), and global (133°) zones. Multivariable logistic regression was used to evaluate associations with metabolic diseases in 4,618 participants. Hypertension showed significant inverse associations with VD and FD across all zones (all P < 0.001). Diabetes exhibited striking regional specificity, characterized by increased TC and MC and decreased FD were confined to the peripheral and global zones (all P< 0.01), with inverse associations for TC and MC detectable only in the global zone (P < 0.05). This study establishes the first DL framework specifically designed for retinal vessel segmentation in true color UWF images and reveals disease-specific regional vascular patterns that would be missed by conventional fundus photography, highlighting the value of UWF imaging for comprehensive systemic disease assessment and providing a biological prior for the development of interpretable and region-aware models. The code will be released on GitHub: https://github.com/yxyXinyue/UWF Retinal-Vasculature-Segmentation-Cardiometabolic.
Xinyue Wang, Xinyue Yang, Yu-Wei Wang et al.· IEEE journal of biomedical a...· 0 citations