TGFusion is proposed, a text-guided latent-space flow matching framework that unifies degradation suppression and cross-modal fusion and achieves superior or competitive performance in perceptual quality, image naturalness, structural-detail preservation, and infrared-saliency retention, while remaining robust across diverse single and compound degradations.
Abstract
Infrared-visible image fusion under realistic degradation scenarios is a challenging task, as degradations not only cause a loss of reliable modality-specific information in observed images but also hinder the fusion process. Recent studies indicate that text can provide prior information about degradation characteristics, complementing the limited evidence available from corrupted input images and facilitating fusion. However, existing methods typically inject fixed global text representations into visual features, making it difficult for textual guidance to adapt to spatially varying degradations, local structures, and thermal saliency. To this end, we propose TGFusion, a text-guided latent-space flow matching framework that unifies degradation suppression and cross-modal fusion. TGFusion encodes task, degradation, and generation cues into structured prompts. To fully exploit these priors, we design a Prompt-conditioned Multi-stream Joint Flow Transformer that represents text as an independent semantic stream alongside fusion, visible, and infrared streams. Joint attention enables token-level bidirectional interaction and layer-wise updating among semantic and visual representations, allowing degradation semantics to dynamically guide reliable information selection and fusion latent generation. Extensive experiments on public benchmarks and complex degradation scenarios demonstrate that TGFusion achieves superior or competitive performance in perceptual quality, image naturalness, structural-detail preservation, and infrared-saliency retention, while remaining robust across diverse single and compound degradations.
Infrared and visible image fusion aims to integrate complementary information from source images to generate high-quality fusion images that serve downstream tasks. However, the differentiated representation of image scene content, the unpredictability of degradation modes in source images, and the complexity of composite degradations pose significant challenges to constructing degradation-robust fusion models. To overcome these challenges, this study proposes STAFuse, a framework designed for composite degradation-robust image fusion that utilizes adaptive degradation-mode identification and aggregated textual prior guidance. We first introduce a scene- and degradation-aware mechanism that extracts crucial context and degradation data, converting it into a textual format to generate dynamic convolution kernels. This allows for the adaptive identification and elimination of varying degradation effects and scene disparities. Additionally, we implement an aggregated scene prior that condenses multi-source information into a fusion text, simulating an ideal scene to effectively guide the retrieval and fusion of multimodal information. Finally, we design a textual-domain supervision loss to perform auxiliary semantic supervision on the fusion model, thereby suppressing degradation effects and improving the visual fidelity of the fusion results. Experimental results demonstrate that the proposed method significantly outperforms existing state-of-the-art methods in terms of robustness, flexibility, and the aggregation of complementary information.
Ting Lv, Hong Jiang, Yu Liu· IEEE Transactions on Image P...· 0 citations
Text-guided image fusion has recently emerged as an effective paradigm for integrating multi-modal information while enabling flexible and task-oriented fusion control. However, existing text-guided fusion methods often rely on shallow semantic-visual interaction and limited attention mechanisms, which restrict their ability to robustly handle complex degradations and fully exploit textual guidance. In this paper, we propose an iterative text-guided image fusion framework that incorporates text-conditioned feature interaction across multiple fusion and refinement stages. The proposed method integrates deepest-level Cross-Attention, multi-scale Cross-Gate Fusion, and stage-specific text-conditioned modulation, allowing the global text embedding to condition hierarchical feature fusion and residual refinement. By repeatedly injecting the pooled text embedding across hierarchical decoder and refinement stages, the proposed framework provides degradation-aware global semantic conditioning while preserving complementary information from the visible and infrared modalities. Experiments on several benchmark datasets show that the proposed method improves several information-preservation and perceptual-quality metrics, while exhibiting metric-dependent trade-offs on some datasets.
Si-Yang Liu, Pei Zhou, Tianhu Jin et al.· 0 citations
Text image super-resolution aims to improve the readability of low-quality text images while preserving character structures, stroke details, and semantic consistency. Compared with natural image super-resolution, this task is more sensitive to structural distortion because small changes in stroke topology may lead to incorrect text recognition. To address this problem, this paper proposes an OCR prior-guided cross-scale framework for text image super-resolution. Specifically, character-level semantic priors extracted from a pretrained OCR model are introduced to provide structural guidance for degraded text reconstruction. A gated feature modulation mechanism is designed to adaptively regulate the contribution of OCR priors, reducing the influence of unreliable semantic predictions. A cross-scale dynamic attention module is also developed to aggregate multi-granularity visual features, enabling the model to jointly recover fine stroke boundaries and global character structures. In addition, a sequence-aware calibration module is introduced to improve structural consistency along the logical reading order of text. Experiments on mixed text image benchmarks and the TextZoom dataset show that the proposed method achieves competitive or better performance among the compared methods in terms of PSNR, SSIM, and recognition-oriented metrics. Additional ablation, OCR prior robustness, and computational complexity analyses further indicate that the proposed framework improves text readability while maintaining a reasonable accuracy–complexity trade-off. The results also suggest that OCR priors are useful for text image reconstruction, but should be used as soft constraints when external recognition predictions are uncertain.
Enhancing degradation robustness is essential for deploying image fusion techniques in real-world dynamic scenes. However, most existing methods either handle a single degradation type or assume fixed multi-degradation settings, making them insufficient for dynamically heterogeneous and composition ally complex degradations in practice. Moreover, they often fail to recover the semantics of salient scene targets when these targets are degraded or missing, leading to weakened semantic representation and reduced target saliency. To address these challenges, we propose DuS-DiFuse, a robust dual-stream latent diffusion framework composed of a diffusion fusion unit and a generative modulation unit. In the diffusion fusion unit, we fine-tune a CLIP visual encoder on multi-source data to perceive degradation types and severities, and employ latent diffusion to uniformly model multi-type, cross-level degradations with varying parameters. A Groupwise Fusion Control Module (GFCM) is further embedded into the latent degradation-removal process, enabling joint modeling of dynamic degradation removal and multimodal information fusion. In the generative modulation unit, pretrained latent diffusion priors are used to remodulate the initial fusion results, enabling controllable semantic restoration and generative enhancement, thereby improving target saliency and overall visual quality. To preserve fine-grained details during latent-to image reconstruction, we introduce a Detail-Restoration Fidelity Module (DRFM), which constrains texture reconstruction by jointly leveraging multi-level skip features from multiple source images and enhances structural fidelity in the fused results. Extensive experiments on multiple fusion datasets demonstrate that DuS-DiFuse achieves leading fusion performance, exhibits strong robustness to heterogeneous degradations, generalizes well across fusion tasks, and supports effective controllable generative modulation.
Lei Cao, Hao Zhang, Peng Zhang et al.· IEEE Transactions on Pattern...· 0 citations
AnchorSteer is proposed, a training-free framework that exerts fine-grained control over both initialization and denoising trajectory that consistently outperforms existing baselines in text--image alignment while preserving high visual quality.
Xinyi Wang, Yuyang Huang, Yalin Su et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.