Interpretable Spectral-State Learning via Unmixing-Guided Mamba Attention Networks for Hyperspectral Image Classification
TL;DR
A spectral-guided state-space attention network (SGSA-Net) for hyperspectral image classification, which incorporates bidirectional Mamba-based spectral sequence modeling, unmixing-aware representation learning, spatial attention, and multi-level explainability is proposed.
Abstract
The classification of hyperspectral images demands models that can capture minor spectral interdependencies and can be interpreted for scientific analysis and operational remote sensing applications. In this work, we propose a spectral-guided state-space attention network (SGSA-Net) for hyperspectral image classification, which incorporates bidirectional Mamba-based spectral sequence modeling, unmixing-aware representation learning, spatial attention, and multi-level explainability. In the proposed architecture, each hyperspectral pixel is represented as an ordered spectral sequence, allowing the bidirectional selective state-space encoder to model input-dependent spectral-state transitions, where the latent state is updated across successive wavelength bands according to the current spectral response and the contextual information accumulated from other spectral regions. We propose an unmixing pretraining stage, initialized using vertex component analysis-based endmembers, to induce physically relevant abundance-aware spectral embeddings prior to supervised fine-tuning. The local spatial spectrum representations are further refined with a depthwise spatial attention module. Complementary interpretability mechanisms are provided by the mamba selection gate, endmember abundances, channel attention, and gradient-based saliency. We performed experiments on three benchmark hyperspectral datasets, namely Indian Pines, Pavia University, and Salinas, which are publicly available from the Kaggle hyperspectral datasets repository. Under a fixed sample-per-class training approach, the SGSA-Net attained overall accuracies of 68.00%, 76.60%, and 92.09% on Indian Pines, PaviaU, and Salinas, respectively, with mean AUC values of 0.961, 0.981, and 0.995, respectively. The results show that the proposed model learns highly discriminative spectral representations, as evidenced by the strong AUC performance across all datasets.