Sparse Tensor Attention for Hyperspectral Unmixing
Abstract
Hyperspectral unmixing (HU) is a fundamental task for resolving the mixed pixel problem by decomposing a hyperspectral image (HSI) into constituent endmembers and their corresponding abundances. Although deep-learning-based HU methods have achieved promising performance, most rely on manually designed network architectures that encode spatial–spectral priors only implicitly, limiting their interpretability and ability to model complex spatial–spectral characteristics. To address this limitation, we propose a self-supervised sparse tensor attention (TA) framework for HU, termed STU. The proposed method builds upon the mathematical connection between self-attention and Tucker decomposition and exploits the nonlocal spatial–spectral self-similarity and low-rank structure inherent in HSIs. Specifically, pairwise dependencies among local spatial–spectral cubes are encoded into a sparse tensorized attention core, which is projected through mode-wise low-rank factor matrices to recover the abundance representations. The resulting abundances are constrained to satisfy the abundance nonnegativity and sum-to-one constraints and are subsequently combined with a learnable endmember matrix under the linear mixture model (LMM) to reconstruct the observed HSI. This formulation integrates data-adaptive nonlocal dependency modeling, multilinear low-rank representation, and physically consistent endmember-abundance estimation within a unified framework. Extensive experiments on synthetic and real-world hyperspectral datasets demonstrate that STU achieves competitive or superior performance compared with representative model-based and deep-learning methods.