Dual-Stream Spatial–Spectral Network with Nested Attention for Hyperspectral Image Classification
Abstract
Hyperspectral image classification (HSI) requires a model to distinguish subtle spectral differences while preserving the spatial structure of land-cover regions. CNN-based methods are effective for local spectral–spatial extraction, but their limited receptive fields can weaken broader context modelling. Transformer-based methods improve long-range dependency modelling, yet fixed patch partitioning may reduce their sensitivity to fine local structures. To address these limitations, this study proposes the Dual-Stream Spatial–Spectral Network with Nested Attention (DSSN), which separates local spectral–spatial feature extraction from multi-scale spatial-context modelling before adaptive fusion. The DSSN combines a cascaded 3D-CNN spectral stream, a nested Transformer spatial stream with pixel-level and patch-level interactions, and a channel-attention-based adaptive fusion module. Experiments on Indian Pines, Pavia University and Salinas show DSSN achieves overall accuracies of 98.11%%, 99.88% and 99.82%, respectively, outperforming other baselines. The ablation experiments confirm that each major component contributes to the final performance. Although the model requires more parameters and longer inference time than several compared baselines, its inference time remains at the millisecond level. These results suggest that decoupled spatial–spectral representation and adaptive multi-scale fusion can improve hyperspectral image classification under the evaluated benchmark settings.