Skip to content
Preprint

Preserve More Details: Mitigating Content Drift in Real-World Image Super-Resolution

Aug 2026 · 0 citations · 53 references
Computer Science

TL;DR

FSP-Diff, a novel one-step diffusion model featuring a dual-pathway architecture that refines semantic guidance using structured details to mitigate semantic deviations, is proposed and demonstrates that FSP-Diff surpasses existing one-step diffusion methods in both quantitative and qualitative metrics.

Abstract

Real-world image super-resolution (Real-ISR) aims to reconstruct high-quality (HQ) images from low-quality (LQ) inputs subject to diverse real-world degradations. Recent advances have leveraged the LQ inputs and natural image priors learned by Stable Diffusion models to achieve impressive results. However, existing methods often overlook insufficient clarity of LQ inputs inevitably induce content drift in the generated HQ images. This manifests primarily as visual detail degradation and textual semantic shift, severely compromising both fidelity and perceptual quality. To address this challenge, we propose FSP-Diff, a novel one-step diffusion model featuring a dual-pathway architecture. This architecture comprises a Detail-Conditioned Pathway for injecting structured details to recover fine structures, and a Detail-Modulated Semantic Pathway that refines semantic guidance using structured details to mitigate semantic deviations. Extensive experiments on standard Real-ISR benchmarks demonstrate that FSP-Diff surpasses existing one-step diffusion methods in both quantitative and qualitative metrics.

View source

Similar papers

Preprint Aug 2026

PixelIR: Fidelity-Perception Decoupling via Pixel-Space Image-Residual Flow Matching for Efficient One-Step Real-World Super-Resolution

Real-world image super-resolution (Real-ISR) aims to preserve structures supported by the degraded observation while reconstructing perceptually realistic details. However, existing Real-ISR methods largely optimize fidelity and perceptual quality within a shared network, causing the two objectives to interfere throughout training and making their balance difficult to control. Recent one-step methods reduce sampling steps, yet often inherit both this coupled optimization behavior and the expensive high-resolution backbone of their multi-step predecessors. We argue that efficient Real-ISR requires not only a shorter sampling trajectory, but also specialized modeling of faithful reconstruction and perceptual detail synthesis. Based on this insight, we propose PixelIR, a fidelity-perception decoupling framework built upon pixel-space image-residual flow matching. PixelIR first learns an image flow that maps the degraded observation to a faithful reconstruction. Then, a residual flow synthesizes the missing perceptual details from noise without repeatedly relearning or overwriting the complete restoration solution. We further distill the teacher into a deployment-oriented one-step student within a coarse-to-fine pyramid architecture. Extensive experiments show that PixelIR achieves leading PSNR, SSIM, and LPIPS on both RealSR and DRealSR. The final model completes pixel-space restoration in a single evaluation with only 32.9M parameters, 89.7G MACs, and 8.5ms latency, demonstrating a strong practical fidelity-perception-efficiency balance.

Bingtian Qiao, Yue Shi, Yong Guo et al. · 0 citations
Preprint Aug 2026

EDITBRIDGE: Towards Faithful and Efficient Ultra-High-Resolution Image Editing

This work proposes EditBridge, a diffusion bridge framework for efficient ultra high-resolution editing that achieves high-fidelity editing with superior perceptual quality at resolutions up to 4K, delivering 3.6--8.4$\times$ speedup at 2K and enabling practical 4K editing in 61 seconds.

Jiayi Song, Shijie Huang, Fangtai Wu et al. · 0 citations
Preprint Aug 2026

When Latents Forget Pixels: Restoring Fidelity in Diffusion Transformer Super-Resolution

This work proposes a pixel-grounded super-resolution (PGSR) framework that preserves LR-observed pixel evidence before VAE compression and reuses it throughout restoration and improves the realism--fidelity trade-off and produces more faithful, visually convincing results than existing latent generative SR approaches.

Yu Shi, Yuyao Zhang, Yu-Wing Tai · 0 citations
Conference Open access Aug 2026

Uncertainty-Guided Latent Diffusion Models for Faithful Super Resolution

UGDiff, a novel diffusion guidance paradigm designed to further improve the perception-distortion balance, is introduced, which first estimates the reconstruction uncertainty of the latent features corresponding to a high-fidelity image and guides the diffusion process to selectively restore high-frequency details in high-uncertainty regions, while preserving fidelity elsewhere.

Ren Wang, Yung-Yu Chuang · 0 citations
Aug 2026

A confidence-guided hybrid network for image restoration

The confidence-guided hybrid network (CGHNet) is proposed, a parallel three-branch framework that jointly performs frequency-decoupled local restoration, global context modeling, and pixel-wise degradation prior estimation and its key component is a confidence-guided feature purification mechanism.

Xiaohui Kou, Yang Yan, Qiuyan Wang et al. · 0 citations
Preprint Aug 2026

PixRestore: Unified Image Restoration via Pixel Diffusion Transformer

PixRestore is presented, a VAE-free pixel-space Diffusion Transformer (DiT) for UIR, where the diffusion backbone is trained entirely from scratch, without relying on T2I pretraining.

Lingchen Sun, Rongyuan Wu, Xiangtao Kong et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.