Depth-Prior Guided Learning for Infrared-Visible Image Registration and Fusion
Infrared and visible image fusion integrates complementary information to enhance scene perception. However, spatial misalignment caused by varying sensor poses and dynamic scenes often degrades the quality of fusion. Most existing approaches focus solely on fusing pre-registered images within the 2-D domain, neglecting the underlying geometric inconsistencies while lacking effective priors to ensure cross-modal structural consistency. To overcome these limitations, we propose a depth-prior guided registration and fusion framework. Firstly, we design a Depth-Guided Cross-Modal Attention (DCMA) module that incorporates depth geometric priors into linear attention, enabling robust and efficient cross-modal feature interaction. Secondly, our DCMA module serves as a unified component across both registration and fusion stages, enabling consistent depth-guided feature interaction for cross-modal alignment and complementary information aggregation. Finally, we introduce a lightweight Depth Quality Assessor (DQA) that generates a continuous quality score to adaptively interpolate between depth-guided and standard attention, maintaining robust performance when depth estimates are unreliable. Comprehensive evaluations on multiple benchmarks demonstrate that our method achieves state-of-the-art performance in both registration and fusion tasks, validating the effectiveness of depth-prior guided cross-modal learning for enhanced scene perception.