Skip to content
Review

Infrared–visible image fusion for robust visual perception: methods, benchmarks, evaluation and open challenges

Aug 2026 · The Visual Computer · Vol 42 · 0 citations · 95 references

TL;DR

This survey reviews recent progress in infrared–visible image fusion, with emphasis on deep learning methods published between 2018 and 2025, and outlines future directions towards robust, interpretable and resource-efficient fusion systems supported by transparent search protocols, community benchmarks, reproducible comparisons and application-specific evaluation.

View source

Similar papers

Open access Sep 2026

A Lighting-Aware Infrared–Visible Image Fusion Network for Security Surveillance

Infrared and visible image fusion combines the thermal cues captured by infrared sensors with the rich structural and texture information provided by visible images. This technique is particularly valuable for security surveillance, nighttime scene perception, and target recognition under challenging environmental conditions. Existing methods generally adopt fixed fusion strategies and neglect dynamic dual-modality changes under low illumination, overexposure and strong light interference, failing to preserve both thermal target saliency and visible structural details in fused images. To address this problem, this paper proposes a Lighting-Aware Spatial–Frequency Fusion Network (LASFNet) for security surveillance. It first estimates modality reliability across low-light, overexposed and infrared-salient regions, incorporating it into the fusion of frequency-domain amplitude and phase. Spatial infrared, visible and frequency-domain compensation features are then jointly fused to generate the output. Experiments on M3FD, MSRS and RoadScene datasets show that LASFNet achieves competitive fusion performance with strong cross-dataset generalization. On M3FD, YOLOv8s with LASFNet-fused inputs achieves 84.825% mAP@0.5 and 57.319% mAP@0.5:0.95, outperforming visible, infrared and YDTR-fused inputs. The proposed method balances target saliency and scene structure, providing more effective visual input for object detection under complex illumination.

Unknown authors · 0 citations
Open access Aug 2026

Residual Conditional Diffusion with Transformer Refinement for Unsupervised Infrared–Visible Image Fusion

Infrared–visible image fusion aims to integrate thermal target information from infrared images and structural texture information from visible images into a single informative image. Existing deep fusion methods still face challenges in preserving fine textures, maintaining structural consistency, and balancing complementary information under low-light conditions. To address these issues, this paper proposes MRCDFusion, an unsupervised infrared–visible image fusion network based on residual conditional diffusion and Transformer refinement. Specifically, a shared dense encoder is used to extract modality-specific and cross-modal complementary features from infrared and visible images. A Modality-Level Attention Module (MLAM) is then introduced to aggregate strong responses from infrared and visible features and construct modality-aware condition features for guiding the diffusion process. Instead of generating fused features from scratch, the proposed method adopts a base-plus-residual diffusion strategy, in which base features preserve global structures and residual diffusion enhances local details. A deterministic noise strategy is further introduced to improve inference reproducibility. The diffusion-enhanced features are refined by a window Transformer and depthwise separable convolutions, followed by gated feature fusion and progressive image reconstruction. Experiments are conducted primarily on the low-light LLVIP dataset, while FMB, TNO, and RoadScene are used for zero-shot cross-dataset evaluation without additional fine-tuning. The results show that MRCDFusion achieves particularly strong performance in gradient- and edge-related metrics while remaining competitive in visual information fidelity and cross-modal correlation metrics. Ablation studies verify the effectiveness of the main components, and downstream detection and auxiliary segmentation experiments further demonstrate the potential utility of the fused representations for subsequent visual perception tasks.

Sirui Huang, Lin Tian, Yao Zhang · 0 citations
Conference Aug 2026

Lightweight infrared and visible image fusion via global selective state-space modeling

Infrared and visible image fusion aims to effectively exploit the strength of the infrared modality in target saliency perception and the complementary capability of the visible modality in representing fine-grained texture details, which is crucial for complex scene understanding and downstream visual tasks. Existing deep learning methods achieve strong performance but often rely on complex cross-modal modeling, leading to large models and high computational costs, limiting real-time deployment. To address these challenges, this paper proposes a lightweight infrared and visible image fusion framework, termed GSSMFuse, based on Global Selective State-Space Modeling. In the feature extraction stage, depthwise separable convolutions are employed to capture local spatial structural information with low computational overhead, followed by a Mamba-based GSSM block to efficiently model long-range dependencies, enabling joint representation of local details and global semantic information. In the fusion stage, infrared and visible features processed by the GSSM block are directly combined and further integrated using lightweight convolution, effectively exploiting the complementary characteristics of the two modalities while maintaining high computational efficiency. Furthermore, a knowledge distillation-based training strategy, together with a structural consistency constraint, is introduced to enhance the fusion quality of the lightweight model by guiding the student network to inherit discriminative representations from the teacher model while preserving a compact architecture. Extensive experiments on the public M3FD dataset demonstrate that the proposed GSSMFuse consistently outperforms existing state-of-the-art fusion methods, while significantly reducing model parameters and computational complexity and achieving competitive performance in downstream object detection tasks.

Wenkuan Xie, Weiguo Pan, Jiancheng Zhang et al. · 0 citations
Aug 2026

Universal Representation for Real-World Misaligned Infrared-Visible Image Fusion.

Infrared and visible image fusion is pivotal for robust visual perception across all weather conditions and scenes. Although deep learning-based methods have made notable progress, most either assume pre-aligned inputs or rely on implicit feature-space alignment, which fails to fundamentally address the amplification of registration errors and the loss of semantic structure in the fused results. To this end, we propose a universal representation and end-to-end framework for jointly registering and fusing unaligned infrared-visible image pairs, dubbed URMIF. Each image is mapped into modality-invariant (homogeneous) and modality-specific (heterogeneous) features: the invariant "structural skeleton" encodes geometry and semantics to stabilize alignment, while the specific "texture carrier" preserves thermal saliency and visible details to enable complementary fusion. Therefore, we propose a bi-directionally coupled registration-fusion module. This module performs hierarchical deformation estimation from coarse to fine, effectively mitigating visual mismatches caused by complex parallax in real-world scenes. Within this framework, the fusion component acts as the "evaluator" of registration, providing feedback regularization to update the deformation and suppress error accumulation. Furthermore, we introduce a dominant-plane prior as a scene-level constraint, seeding stable global and patch-wise homographies and reconciling cross-modal detail conflicts, to reinforce geometric consistency and semantic reliability. We also release a large-scale dataset comprising 1,500+ unaligned infrared/visible pairs with registration ground truth, spanning diverse illumination conditions and fields of view. Based on this dataset and additional benchmarks, extensive experiments validate that our framework achieves robust alignment and high-quality fusion on misaligned inputs, markedly reducing artifacts and improving the performance of downstream tasks such as detection and segmentation. Code and benchmark are available at https://github.com/ZengxiZhang/URMIF.

Jinyuan Liu, Zengxi Zhang, Jiahao Zhang et al. · 0 citations
Open access Jul 2026

AQFS-Net: An Adaptive Quality-Aware Fusion and Saliency-Guided Network for Visible-Infrared Object Detection

Object detection in real-world scenarios is often challenged by adverse visual conditions, such as low illumination, strong glare, and dense fog, which severely degrade visible-spectrum features and lead to missed detections, inaccurate localization, and reduced detection accuracy. To address these issues, this paper proposes AQFS-Net, a dual-modal fusion detection network for visible-infrared object detection. Built upon YOLOv13, AQFS-Net adopts a symmetric dual-branch backbone by incorporating infrared images, thereby exploiting the complementary information between the visible and infrared modalities. To alleviate the negative transfer caused by conventional static fusion strategies, an Adaptive Quality-Aware Fusion Module (AQFM) is designed to dynamically enhance informative features and suppress degraded information according to the modality-specific reliability of different regions. In addition, a Foreground-Aware Saliency Guidance (FASG) branch is introduced to guide the network to focus on target regions through foreground supervision, reducing interference from complex backgrounds. Experimental results on the public LLVIP and M3FD datasets show that the proposed method improves mAP@0.5 by 6.8 and 3.1 percentage points, respectively, compared with the baseline using only visible images. These results demonstrate the effectiveness of AQFS-Net in improving dual-modal fusion quality and detection performance under challenging visual conditions, providing a practical reference for visible-infrared object detection in complex illumination scenarios.

Weijun Wu, Xufei Zhuang · 0 citations
Open access Aug 2026

Infrared and Visible Image Fusion via Style-Based Recalibration and Edge Enhancement

Infrared and visible image fusion (IVIF) aims to preserve infrared thermal targets and visible structural textures in one informative image. Although recent attention-based methods improve cross-modal interaction, their post-fusion refinement remains limited in two aspects: modality-specific channel statistics are no longer explicitly exposed after feature mixing, and repeated attention-based aggregation can smooth spatial responses and weaken high-frequency visible details. To address these issues, this work proposes a lightweight end-to-end IVIF network with two complementary refinement modules. MSG carries out cross-modal style-based recalibration by making use of the joint mean and standard deviation of the two pre-fusion encoder features, so that first- and second-order pre-fusion modality statistics can guide post-fusion channel selection. DGM carries out edge enhancement by constructing a parameter-free Sobel detail prior from source images and learning only a lightweight residual modulation to perform restoration of high-frequency evidence. With only 80,160 trainable parameters, the proposed method achieves the best or tied-best value on three of seven standard fusion-quality metrics on FMB and four of seven on LLVIP, and ablation results further confirm the complementary effects of MSG and DGM.

Wenhua Zhao, Lei Zhong · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.