Self-Supervised Hyperspectral Super-Resolution From RGB With Score-Distilled Spectral Prior
Recovering hyperspectral images (HSIs) from RGB observations is a highly ill-posed problem due to severe spectral information loss. However, current methods either rely on costly paired RGB–HSI datasets that are difficult to obtain or on unpaired RGB–HSI data for spectral guidance, which increases training costs and often leads to physically distorted spectral reconstructions. To address these challenges, we propose a self-supervised score-distilled spectral prior (SDSP) framework, which consists of a wavelet-based cross-attention reconstruction network (WCAR-Net) as the generator and a pretrained spectral diffusion model as the data-driven spectral expert. Score distillation sampling (SDS) is applied to the diffusion model to compute spectral gradients, which are incorporated into the generator’s loss to iteratively refine its predictions and produce high-quality hyperspectral reconstructions. In WCAR-Net, a cross-attention refinement module is designed to fuse spatial details with high-level semantic features, while wavelet-based feature decomposition preserves fine frequency-domain structures. Guided by the spectral expert provided by the diffusion model, the network is trained in a self-supervised and iterative manner, thereby eliminating the need for paired RGB–HSI data. Extensive experiments on multiple datasets demonstrate that our method significantly outperforms state-of-the-art approaches in both spectral fidelity and spatial reconstruction.