Skip to content

MTVDiff: Multimodal Conditional Latent Diffusion for Enhanced Thermal-to-Visible Face Translation

Jul 2026 · arXiv.org · Vol abs/2607.19886 · 0 citations · 30 references
Computer Science

TL;DR

This work provides a robust solution for face recognition systems operating under varying illumination conditions and advances the state-of-the-art in cross-spectral facial image translation through effective multimodal integration.

Abstract

Thermal-to-visible face translation presents fundamental challenges including geometric discontinuities, semantic attribute mismatches, and identity degradation. We propose MTVDiff, a novel multimodal latent diffusion framework that synergistically integrates depth and textual information to address these limitations while preserving identity characteristics. The MTVDiff framework presents three core technical contributions: (1) a Dual-Branch Cross-Attention Fusion (DBCAF) module for multi-scale thermal-depth feature extraction and fusion; (2) a Gated Text-to-Visual Feature Alignment mechanism for semantically-guided generation; and (3) Spatial Feature Transformations (SFT) for adaptive multimodal prior integration. Extensive experiments on the MCXFace and SpeakingFaces datasets demonstrate that our multimodal approach significantly outperforms existing GAN-based and diffusion-based approaches across multiple metrics, achieving substantial improvements in both image quality and face verification performance, with FID reductions of up to 48.3% and Rank-1 accuracy improvements of up to 8.9\%. Our work provides a robust solution for face recognition systems operating under varying illumination conditions and advances the state-of-the-art in cross-spectral facial image translation through effective multimodal integration.

View source

Similar papers

Jul 2026

Multi-condition guided diffusion model for face sketch-to-photo synthesis.

A diffusion-based framework with a stage-wise multi-condition guidance mechanism that enhances both structural and textural fidelity and compares with recent image-to-image translation and diffusion-based baselines to observe competitive performance in both visual coherence and identity preservation.

Yue Que, Xuegui Cheng, Shuqian Shi et al. · 0 citations
Open access Aug 2026

Edge-Oriented Lightweight YOLO-World with Cross-Modal Fusion Adaptation for Multimodal Object Detection and Deployment

Multimodal detection models enable flexible object detection through text prompts, but YOLO-World-style models still incur high computational and storage costs on edge devices. To address this problem, this paper develops an edge-oriented adaptation framework for replacing the original visual encoder of YOLOv8l-worldv2 with a compact visual branch. The framework jointly considers visual-branch compression, multi-scale interface consistency, cross-modal feature compatibility, and edge-side inference, rather than optimizing these aspects independently. Specifically, a YOLOv7-based visual branch is reconstructed using depthwise separable convolutions, enhanced by DyHead, and compressed through sensitivity-guided Filter Pruning via Geometric Median (FPGM) under multi-scale interface constraints. An identity-initialized semantic adaptation layer and multi-template text prototypes are then introduced to alleviate the feature distribution mismatch between the compressed visual branch and the original cross-modal fusion space. Finally, offline text prototype generation and TensorRT-based INT8/FP16 mixed-precision inference are used for edge deployment. Experiments on BDD100K show that the lightweight visual encoder achieves 62.45% mAP@0.5 with 22.5 M parameters and 52.8 GFLOPs. After multimodal integration, the proposed model achieves 64.9% mAP@0.5 with 31.5 M parameters and 67.2 GFLOPs, reducing GFLOPs by 67.1% compared with YOLOv8l-worldv2 while causing only a 1.6 percentage-point accuracy drop. On the Jetson Orin Nano Super, the deployed model reaches 32.6 FPS with a model size of 32 MB, demonstrating its feasibility for edge perception scenarios.

Mengnan Jiang, Tian-Li Mo, Jie Hu et al. · 0 citations
Preprint Sep 2026

Radiation, Rotation and Scale Invariant Feature Descriptor for Multimodal Image Matching

Multimodal image matching is a fundamental task for multi-source information fusion. However, geometric distortions and nonlinear radiometric differences (NRD) severely limit performance, especially under radiometric, rotation, and scale variations. To address this issue, we propose a radiation, rotation, and scale invariant (RRSI) feature descriptor. First, a dual-head regional sampling (DHRS) module simultaneously performs Cartesian and Log-Polar sampling on keypoint neighborhoods, retaining spatial structural properties while enhancing robustness to rotation and scale variations. We then jointly encode geometric and radiometric relations between multimodal images in a unified deep feature space, enabling feature encoding, interaction, and fusion across intra-modal, dual-head sampled, and inter-modal regions. Furthermore, we introduce a bidirectional cross-modal generative reconstruction constraint during training. By decoding implicit features into structural patches of the counterpart modality, this mechanism anchors modality-invariant geometric topologies without additional inference overhead. Experiments on optical-infrared and optical-SAR datasets demonstrate highly competitive matching performance and strong robustness to rotation and scale variations. RRSI supports the full rotation range from 0 to 360 degrees and scale factors up to four. Its generalization ability is further validated on multimodal images from computer vision, remote sensing, and medical imaging. The implementation will be made publicly available at https://github.com/yeyuanxin110/RRSI .

Unknown authors · 0 citations
Preprint Aug 2026

Unleashing the Power of Text: Text-Guided Flow Matching for Image Fusion under Complex Degradations

TGFusion is proposed, a text-guided latent-space flow matching framework that unifies degradation suppression and cross-modal fusion and achieves superior or competitive performance in perceptual quality, image naturalness, structural-detail preservation, and infrared-saliency retention, while remaining robust across diverse single and compound degradations.

Axi Niu, Jiehua Li, Kang Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.