Skip to content

Decoupling Cross-Modality Manifold Discrepancy: Leveraging Visible Diffusion Priors for Infrared Super-Resolution

Jul 2026 · arXiv.org · Vol abs/2607.21174 · 0 citations · 55 references
Computer Science

TL;DR

A dual-path diffusion-based framework for IISR with a Global Representation Modulation module to extract modality-specific information from infrared imagery and guide the global distribution of the diffusion model toward the ground truth is proposed.

Abstract

Infrared image super-resolution (IISR) mitigates the limitations imposed by low spatial resolution. Existing methods have recognized that IISR should preserve consistency in global distribution and structural information while enhancing image clarity. However, these methods are either insufficient or overly intrusive, a problem that becomes even more pronounced in diffusion-based models. To address these issues, we propose a dual-path diffusion-based framework for IISR, termed Shift-IISR. The proposed method is designed to improve the consistency of IISR results while preserving the generative capacity of diffusion models. Specifically, we develop a Global Representation Modulation (GRM) module to extract modality-specific information from infrared imagery and guide the global distribution of the diffusion model toward the ground truth. In addition, we introduce a Local Structure Refinement (LSR) module to encourage the model to focus on structural information at each step of the iterative denoising process. Extensive experiments demonstrate that the proposed method effectively improves distributional and structural consistency while maintaining competitive super-resolution performance. The source code of the proposed Shift-IISR can be available at https://github.com/Assassink8/Shift-IISR.

View source

Similar papers

Open access Aug 2026

A Hybrid Prior-Based Framework for Infrared Image Enhancement Towards Reliable Scene Interpretation

A Hybrid Prior Enhanced Decomposition (HPED) model is proposed, a training-free, model-driven framework that incorporates structural and luminance priors into a multi-stage enhancement pipeline, outperforming all competing methods and suggesting potential applicability in downstream machine perception tasks such as detection and tracking.

Jie Li, Cheng Wang, Xiang-Yu Li et al. · 0 citations
Aug 2026

Rethinking Multi-modal Image Super-resolution: The Key Role of Cross-modal Consistency Prior.

For multi-modal image super-resolution (MISR), exploring cross-modal consistency is of vital importance. However, most existing consistency priors struggle to preserve high-frequency components and fail to provide generalizable regularization, often resulting in blurred edges or inaccurate textures. In this paper, we revisit the pivotal role of consistency prior in MISR task and present an important finding: the modality gap in the Laplacian response between guidance and target images does not conform to the widely adopted Gaussian or Laplacian distributions. Instead, it aligns better with T-distribution. Based on this insight, we propose a T-distribution formed Laplacian response Consistency (TLC) model. This model integrates a T-distribution based Multi-modal Consistency (TMC) prior with a learnable regularization term that operates between the guidance and target images. Additionally, we introduce a Multiplicative Degradation (MD) matrix to model the degradation process from high-resolution (HR) to low-resolution (LR) target images, thereby enabling adaptive non-uniform degradation modeling. The iterative optimization steps of the TLC model are subsequently unfolded into an interpretable network, termed TLCNet. The performance of TLCNet is evaluated on nine datasets across three MISR tasks, demonstrating its superior super-resolution performance compared to other state-of-the-art approaches. The visualization of intermediate features and the causal analysis of the guidance image further confirm its good interpretability.

Jingyi Xu, Xin Deng, Yutong Wang et al. · 0 citations
#machine learning Preprint Sep 2026

Perceptually Regularized Diffusion Model for Image Super-Resolution

Image super-resolution, which aims to reconstruct high-resolution images from their low-resolution observations, is fundamental to medical imaging, remote sensing, surveillance, microscopy, and scientific visualization. Traditional model-based methods formulate super-resolution as an inverse problem with hand-crafted regularization priors. While interpretable and theoretically grounded, they rely on fixed assumptions and require computationally intensive iterative solvers. Deep learning methods offer data-driven flexibility by learning nonlinear mappings from low- to high-resolution images, among which diffusion models have achieved particularly impressive perceptual quality. However, the standard diffusion training objective is a pixel-domain noise-prediction loss that does not explicitly enforce perceptual fidelity, which can lead to oversmoothing and loss of fine image structure. To address these limitations, we propose a perceptually regularized diffusion framework that incorporates prior knowledge through perceptual-loss-based regularization, improving training convergence and encouraging the recovery of meaningful image features. Experiments on benchmark datasets demonstrate improved perceptual quality and competitive distortion metrics, highlighting the effectiveness of regularization for diffusion-based super resolution.

Chuxiang Wang, Pavithra Venkatachalapathy, Ying Liang et al. · 0 citations
Jul 2026

BeyondFusion: Self-Aligned Latent Diffusion for Calibration-Free Infrared Super-Resolution and Infrared-Visible Fusion

Mobile infrared-visible imaging typically pairs a compact infrared sensor with a high-resolution visible camera for complementary perception. While cross-sensor misalignment caused by different optics, viewpoints, fields of view, and exposure timings hinders practical deployment. In this paper, we propose BeyondFusion, a unified latent diffusion framework for calibration-free visible-guided infrared super-resolution and infrared-visible fusion tasks. The proposed framework supports both task-specific training and joint training where two tasks are optimized and executed as two readouts of the same generative process. Instead of relying on explicit registration or geometric warping, BeyondFusion introduces a cross-modal self-aligning (CMSA) module into the denoising U-Net. CMSA reorganizes infrared and visible latent tokens into a shared attention space to learn content-adaptive cross-modal correspondence during the denoising process. Together with misalignment augmentation module, the model is facilitated to exploit visible structural and semantic cues while preserving thermal consistency, enabling high-frequency infrared reconstruction and informative fused-image generation under uncalibrated conditions. Extensive experiments on public benchmarks and a mobile infrared-visible imaging system show strong performance across aligned inputs, low-resolution infrared observations, synthetic misalignments, and real mobile captures with unsynchronized sensors. Ablation studies, unified training analysis, and downstream pedestrian detection further validate the effectiveness of BeyondFusion for calibration-free multimodal imaging.

Minchong Chen, Xiaoyun Yuan, Minyu Cao et al. · 0 citations
Preprint Aug 2026

Preserve More Details: Mitigating Content Drift in Real-World Image Super-Resolution

FSP-Diff, a novel one-step diffusion model featuring a dual-pathway architecture that refines semantic guidance using structured details to mitigate semantic deviations, is proposed and demonstrates that FSP-Diff surpasses existing one-step diffusion methods in both quantitative and qualitative metrics.

Chunxiao Liu, Wei Liu, Anbin Xiong et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.