Sep 2026· Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence· 0 citations· 46 references
TL;DR
AIR-Fusion is proposed, a parameter-efficient adaptation of a frozen, restoration-capable latent diffusion backbone for degraded IVIF without full fine-tuning, transferring restoration priors for robust fusion.
Abstract
Supervised infrared-visible image fusion (IVIF) often overfits limited training distributions, creating a critical generalization gap under open-world degradations (rain, haze, low light, noise, blur). To address this issue, we propose AIR-Fusion, a parameter-efficient adaptation of a frozen, restoration-capable latent diffusion backbone for degraded IVIF without full fine-tuning, transferring restoration priors for robust fusion. A Cross-Modal Bridging Adapter (CMBA) aligns infrared cues and textual instructions with the frozen diffusion conditioning space and injects them into multi-scale denoising features to steer instruction-guided restoration-aware fusion. In addition, a Trajectory-Constrained Rectifier (TCR) regularizes stochastic sampling via a pixel-latent closed loop, rectifying intermediate predictions with source-referenced structures and re-encoding them to stabilize the denoising trajectory and recover fine details suppressed by latent compression. Experiments across multiple datasets and degradation settings show consistent improvements in restoration quality and fusion fidelity, with strong generalization under complex and compounded degradations.
PixRestore is presented, a VAE-free pixel-space Diffusion Transformer (DiT) for UIR, where the diffusion backbone is trained entirely from scratch, without relying on T2I pretraining.
Ling-Chen Sun, Rong-Yuan Wu, Xiang-Tao Kong et al.· 1 citation
MGN-AIR is presented, a novel pixel-level restoration framework for all-in-one image restoration that leverages both textual and visual prompts to provide global and local degradation cues, guiding the model on where to look and how to restore at each pixel.
Chun-Xiao Liu, Wei Liu, Anbin Xiong et al.· 0 citations
DARD first extracts image-specific physical priors from the degraded input through a test-time degradation-aware Retinex decomposition, thereby providing reliable structural guidance for zero-shot restoration and injects these priors into reverse diffusion through a timestep-adaptive frequency fusion strategy to balanc...
Wen-Jie Cai, Yuezhe Yang, Jian-Yang Xia et al.· IEEE transactions on circuit...· 0 citations
A lightweight end-to-end IVIF network with two complementary refinement modules that achieves the best or tied-best value on three of seven standard fusion-quality metrics on FMB and four of seven on LLVIP, and ablation results further confirm the complementary effects of MSG and DGM.
Wenhua Zhao, Lei Zhong· Applied Sciences· 0 citations
This work proposes a pixel-grounded super-resolution (PGSR) framework that preserves LR-observed pixel evidence before VAE compression and reuses it throughout restoration and improves the realism--fidelity trade-off and produces more faithful, visually convincing results than existing latent generative SR approaches.
Single-pixel imaging (SPI) is strongly affected by a hybrid Poisson–Gaussian noise. Clean labels are difficult to obtain in many SPI experiments, while conventional label-free denoisers generally operate after reconstruction and do not explicitly use the bucket-measurement physics. To address this issue, we propose a s...
Xiao-Fei Zhang, Hao Tang, Shu-Run Wang et al.· IEEE Sensors Journal· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.