Jul 2026· IEEE Transactions on Image Processing· Vol 35, pp. 8370-8385· 0 citations· 86 references
MedicineComputer Science
TL;DR
A Degradation-Aware Cross-Modal Prompt Compensation Network that leverages cross-modal degradation cues from a pre-trained vision-language model to condition restoration features within a unified backbone, achieving robust and perceptually consistent restoration across diverse weather conditions.
Abstract
Adverse weather causes diverse and complex image degradations, severely compromising the reliability of computer vision systems. Existing all-in-one restoration models attempt to address multiple degradation types within a unified framework, but often lack explicit spatial and semantic modeling of degradation characteristics, limiting their adaptability to diverse weather conditions. To address this limitation, we propose a Degradation-Aware Cross-Modal Prompt Compensation Network (DCMPC-Net) that leverages cross-modal degradation cues from a pre-trained vision-language model to condition restoration features within a unified backbone. Specifically, our DCMPC-Net mainly consists of the Cross-Modal Prompt Generator (CMPG), Prompt-Guided Attention Alignment Module (PGAAM), and Dual Feature Compensation Module (DFCM). The CMPG integrates textual embeddings with visual features to produce degradation-aware prompts that encode degradation-related semantic and contextual cues. These prompts are injected into the decoder via a PGAAM, which adaptively aligns semantic information with degraded regions to facilitate context-aware restoration. To further enhance structural fidelity, DFCM is introduced that disentangles degradation artifacts from scene structures, thereby improving the reconstruction of fine textures and detailed content. By integrating cross-modal semantic guidance with spatial alignment and structural enhancement, DCMPC-Net achieves robust and perceptually consistent restoration across diverse weather conditions. Extensive experiments show that DCMPC-Net outperforms state-of-the-art methods in both task-specific and unified settings, achieving superior accuracy and visual fidelity. The code is available at https://github.com/fanamber831/DCMPC-Net
TGFusion is proposed, a text-guided latent-space flow matching framework that unifies degradation suppression and cross-modal fusion and achieves superior or competitive performance in perceptual quality, image naturalness, structural-detail preservation, and infrared-saliency retention, while remaining robust across diverse single and compound degradations.
Axi Niu, Jiehua Li, Kang Zhang et al.· 0 citations
Restoring images degraded by adverse weather remains challenging due to spatially heterogeneous degradations. Many existing weather-specific restoration models rely on weather-agnostic global aggregation, naive cross-scale fusion, and deterministic objectives, which struggle to handle heterogeneous degradations in all-in-one adverse-weather settings. To address these limitations, we propose an Uncertainty-guided Adverse-weather Restoration Network (UAR-Net), a weather-specific AiO framework that integrates a gated transformer with balanced multi-scale skip connections. Specifically, we employ Gated Dual-scale Transformer Blocks (GDTB) to jointly model selective global interactions and multi-scale local structures, a progressive Balanced Multi-scale Skip Connection (BMSC) for balanced multi-scale feature integration, and an Uncertainty-Aware Refinement Head (URH) that performs artifact removal, detail enhancement, and predictive uncertainty estimation. The model is supervised by a Brightness-Aware Energy Loss (BAE-Loss) to encourage accurate reconstruction with well-calibrated uncertainty. Extensive experiments demonstrate that our method achieves state-of-the-art performance across multiple adverse-weather benchmarks. The codes will open source upon acceptance.
Zhe-Ke Jin, Yuning Cui, Tianhu Jin et al.· 0 citations
The core of SSVAL is Visual Anchor Prompt Injection (VAPI), which introduces prompts that absorb rich knowledge from external VFMs during training, enabling them to serve as stable visual anchors that mitigate representation deviation during inference.
Qian-Long Yang, Bowen Ye, Xianda Guo et al.· 0 citations
Adverse weather image restoration aims to recover clean background scenes from images degraded by various weather conditions, such as haze, rain, and snow. With the rapid development of deep learning, single-task restoration methods targeting specific weather types have achieved remarkable progress and attracted increasing attention in recent years. More recently, to address the limited generalization of task-specific models, All-in-One (AiO) methods have emerged to handle multiple degradations within a unified framework. However, existing surveys mostly focus on individual degradation types or specific restoration paradigms, and unified reviews of deep learning-based adverse weather restoration are still limited. In this paper, we present a comprehensive survey that jointly organizes single-task and AiO restoration models from the perspectives of network architectures and learning paradigms. We further review widely used datasets, loss functions, and evaluation metrics across different restoration tasks. In addition, we summarize benchmark results of representative methods on public datasets to analyze their performance and generalization ability. Finally, we discuss key challenges and promising research directions to support future developments in this rapidly evolving field.
Enhancing degradation robustness is essential for deploying image fusion techniques in real-world dynamic scenes. However, most existing methods either handle a single degradation type or assume fixed multi-degradation settings, making them insufficient for dynamically heterogeneous and composition ally complex degradations in practice. Moreover, they often fail to recover the semantics of salient scene targets when these targets are degraded or missing, leading to weakened semantic representation and reduced target saliency. To address these challenges, we propose DuS-DiFuse, a robust dual-stream latent diffusion framework composed of a diffusion fusion unit and a generative modulation unit. In the diffusion fusion unit, we fine-tune a CLIP visual encoder on multi-source data to perceive degradation types and severities, and employ latent diffusion to uniformly model multi-type, cross-level degradations with varying parameters. A Groupwise Fusion Control Module (GFCM) is further embedded into the latent degradation-removal process, enabling joint modeling of dynamic degradation removal and multimodal information fusion. In the generative modulation unit, pretrained latent diffusion priors are used to remodulate the initial fusion results, enabling controllable semantic restoration and generative enhancement, thereby improving target saliency and overall visual quality. To preserve fine-grained details during latent-to image reconstruction, we introduce a Detail-Restoration Fidelity Module (DRFM), which constrains texture reconstruction by jointly leveraging multi-level skip features from multiple source images and enhances structural fidelity in the fused results. Extensive experiments on multiple fusion datasets demonstrate that DuS-DiFuse achieves leading fusion performance, exhibits strong robustness to heterogeneous degradations, generalizes well across fusion tasks, and supports effective controllable generative modulation.
Lei Cao, Hao Zhang, Peng Zhang et al.· IEEE Transactions on Pattern...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.