Rethinking Multi-modal Image Super-resolution: The Key Role of Cross-modal Consistency Prior.
Abstract
For multi-modal image super-resolution (MISR), exploring cross-modal consistency is of vital importance. However, most existing consistency priors struggle to preserve high-frequency components and fail to provide generalizable regularization, often resulting in blurred edges or inaccurate textures. In this paper, we revisit the pivotal role of consistency prior in MISR task and present an important finding: the modality gap in the Laplacian response between guidance and target images does not conform to the widely adopted Gaussian or Laplacian distributions. Instead, it aligns better with T-distribution. Based on this insight, we propose a T-distribution formed Laplacian response Consistency (TLC) model. This model integrates a T-distribution based Multi-modal Consistency (TMC) prior with a learnable regularization term that operates between the guidance and target images. Additionally, we introduce a Multiplicative Degradation (MD) matrix to model the degradation process from high-resolution (HR) to low-resolution (LR) target images, thereby enabling adaptive non-uniform degradation modeling. The iterative optimization steps of the TLC model are subsequently unfolded into an interpretable network, termed TLCNet. The performance of TLCNet is evaluated on nine datasets across three MISR tasks, demonstrating its superior super-resolution performance compared to other state-of-the-art approaches. The visualization of intermediate features and the causal analysis of the guidance image further confirm its good interpretability.