Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Aug 2026

Rethinking Multi-modal Image Super-resolution: The Key Role of Cross-modal Consistency Prior.

For multi-modal image super-resolution (MISR), exploring cross-modal consistency is of vital importance. However, most existing consistency priors struggle to preserve high-frequency components and fail to provide generalizable regularization, often resulting in blurred edges or inaccurate textures. In this paper, we revisit the pivotal role of consistency prior in MISR task and present an important finding: the modality gap in the Laplacian response between guidance and target images does not conform to the widely adopted Gaussian or Laplacian distributions. Instead, it aligns better with T-distribution. Based on this insight, we propose a T-distribution formed Laplacian response Consistency (TLC) model. This model integrates a T-distribution based Multi-modal Consistency (TMC) prior with a learnable regularization term that operates between the guidance and target images. Additionally, we introduce a Multiplicative Degradation (MD) matrix to model the degradation process from high-resolution (HR) to low-resolution (LR) target images, thereby enabling adaptive non-uniform degradation modeling. The iterative optimization steps of the TLC model are subsequently unfolded into an interpretable network, termed TLCNet. The performance of TLCNet is evaluated on nine datasets across three MISR tasks, demonstrating its superior super-resolution performance compared to other state-of-the-art approaches. The visualization of intermediate features and the causal analysis of the guidance image further confirm its good interpretability.

Jingyi Xu, Xin Deng, Yutong Wang et al. · 0 citations
Aug 2026

P4VC: Positive Perturbation based Perceptual Preprocessing Framework for Video Compression.

Recent advancements in deep learning have significantly propelled the enhancement of video compression frameworks, encompassing both encoder-side and postprocessing methods. However, these extensively explored methodologies have reached their limits, offering diminishing returns for further improvement. To overcome the constraints of the above optimization patterns, we propose to step beyond conventional frameworks and focus on preprocessing ahead of compression. For preprocessing, the black-box nature of video codecs introduces challenges for deep learning-based optimization: 1) the absence of ground-truth preprocessed videos, and 2) the lack of end-to-end optimization mechanism. To address these challenges, we introduce P4VC, a Positive Perturbation based Perceptual Preprocessing framework that generates adaptive perturbations before compression to enhance the rate-perception trade-off. Specifically, P4VC develops an alternative updating optimization scheme with 1) individual optimization phase that employs an attack-based method to generate multiple codec-friendly positive perturbations for each training sample, directly targeting the practical codec, and 2) universal optimization phase that trains a lightweight preprocessing network to generalize across arbitrary videos, guided by the positive perturbations and developed distribution-aware adversarial learning scheme. Extensive experiments across five codecs, two datasets and six perceptual metrics, demonstrate that P4VC consistently achieves significant compression gains, superior generalization, and real-time preprocessing at 241 FPS.

Mai Xu, Yichen Guo, Shang-Mou Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.