Diffusion-based methods have achieved impressive performance in real-world image super-resolution (Real-ISR) by leveraging large pre-trained stable diffusion (SD) models as powerful generative priors. However, these methods still face two key limitations. First, existing SD-based one-step and multi-step Real-ISR approaches adopt a unified processing paradigm for all input samples, ignoring the varying restoration difficulty across images. Second, the aggressive resolution reduction of the VAE in SD models (e.g., 8x downsampling) leads to irreversible loss of fine-scale details, which cannot be recovered by the subsequent diffusion process. To address these limitations, we propose a Difficulty-aware Dynamic Routing (DDR) strategy that overcomes the rigid, one-size-fits-all processing paradigm. Specifically, we first design a difficulty estimator to predict the restoration cost of each input image, enabling automatic assignment to a network of appropriate capacity. Then, we construct a set of Real-ISR networks with varying model capacities by modulating the spatial downsampling ratio of the VAE in the SD backbone, thereby preserving more high-frequency information for challenging cases while maintaining efficiency for simpler inputs. Extensive experiments have demonstrated the superior efficiency and effectiveness of the proposed model compared to recent state-of-the-art methods.
Xue Wu, Kang Zhao, Kafeng Wang et al.· arXiv.org· 0 citations
The purpose of face enhancement tasks is to improve the recognition of faces, thus adapting to diverse visualization and recognition demands. However, the performance of the majority methods is drastically degraded under extreme conditions, including large pose variations, low resolution, blur, occlusion, and illumination changes, which can distort facial geometry and identity related details. In this work, we construct a simple and effective face robust enhancement method. In particular, in order to maintain the identity consistency of the reconstructed face, an evolutionary learning framework for face disentanglement representation is proposed, in which we disentangle the identity and pose information of the face and unite it with identity recognition as a multi-objective optimization problem, where reconstruction, adversarial, and identity-preserving objectives are adaptively balanced. Further, in order to maintain the pose consistency of reconstructed faces, we construct a unified face pose dictionary, which forms a robust and standard pose representation by statistics and induction of the geometric structure of a large number of face images. In the conditional generation architecture, the pose dictionary could accurately guide the model to realize face reconstruction with desired poses. Extensive benchmark experiments on MS1M, LFW, CPLFW, CFP-FF, CFP-FP, and AgeDB show that the proposed method not only significantly outperforms state-of-the-art methods, but also can further stimulate the discrimination potential of existing face recognition models. Specifically, DiEL achieves an average improvement of 4.66% over the SOTA methods across six benchmark datasets, with particularly significant gains on challenging cross-pose benchmarks such as CPLFW and CFP-FP.
Jingwei Xin, Tian Yang, Jun Hao et al.· IEEE Transactions on Informa...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.