A unified generative adversarial network (GAN) framework that integrates three novel, complementary mechanisms that generates high-fidelity, structurally faithful results from a single reference image, offering a robust solution for applications in virtual try-on, animation, and human image synthesis under challenging pose transformations.
Abstract
Pose transfer, a core task in human-centric image generation, aims to synthesise photorealistic images of a subject in novel poses while preserving identity and intricate clothing details. Existing methods, particularly under large pose variations such as extreme articulation or self-occlusion, often struggle with preserving fine-grained textures and maintaining structural consistency, leading to artifacts like distorted limbs and lost details. To address these challenges, we introduce a unified generative adversarial network (GAN) framework that integrates three novel, complementary mechanisms. First, a Hierarchical Semantic Aligner (HSA) establishes multi-scale semantic correspondence between source appearance and target pose features through local attention and global gating. Second, a Pose-Aware Feature Injection (PAFI) module explicitly models source-target pose discrepancy to generate dynamic modulation parameters for adaptive feature adjustment during decoding. Third, a Gated Residual Fusion (GRF) strategy adaptively balances local detail and global structural information via a learnable dual-branch gating mechanism. Evaluated on the DeepFashion dataset, our framework demonstrates significant improvements, achieving a 13.9% reduction in Fréchet Inception Distance (FID) compared to the MAGPT method, alongside superior scores in Structural Similarity Index (SSIM) and Learned Perceptual Image Patch Similarity (LPIPS). Ablation studies confirm the individual contributions of each component. The proposed approach generates high-fidelity, structurally faithful results from a single reference image, offering a robust solution for applications in virtual try-on, animation, and human image synthesis under challenging pose transformations.
A feature-space alignment framework for person-to-person multi-pose virtual try-on, which estimates appearance flow in a latent feature space and uses the learned flow to warp body-part and garment-related representations to improve realism and structural fidelity.
Yang Yang, Erwei Yin· The Visual Computer· 0 citations
Pose-guided human image generation aims to synthesize an image of a target person based on a reference image and a target pose. Although diffusion-based methods have recently achieved significant progress in pose alignment and visual realism, generating high-fidelity images in regions involving complex pose interactions remains a challenge. Typical issues include structural confusion of limbs and texture distortion in occluded areas. To address this challenge, this paper proposes the HiPA-Gen framework, which is designed to automatically perceive complex interaction regions and generate high-fidelity content within them. Unlike existing methods that mainly rely on flat skeleton conditions or implicit attention responses, HiPA-Gen explicitly models hierarchical spatial relations and local appearance priors in complex interaction regions. Specifically, we design a Dual-Agent Hierarchical Reasoning Module (DHR) where a Prompt-Guided Reasoning Agent identifies interaction-related body parts and a Hierarchical Rendering Agent converts the inferred relations into a Hierarchical Correspondence Pose (HCP) map. The HCP map provides explicit front–back structural cues for limb overlap, hand occlusion, and body self-occlusion. To further reduce local texture ambiguity, we introduce a Part Prior Detail Alignment Module (PDA), which extracts Regional Reference Assets (RRAs) from the source image under HCP-guided part priors. These regional references are then injected into the Pose-Conditioned Interaction-Aware Detail Synthesis Network (PDS) through a Local Enhancement Branch (LEB), enabling more accurate local feature fusion during diffusion-based synthesis. Experiments on DeepFashion and Market-1501 show consistent numerical improvements under the reported evaluation protocols in structural similarity, perceptual quality, and distributional fidelity. Qualitative results further show that the proposed framework produces clearer limb boundaries, more reliable spatial ordering, and more consistent clothing textures in challenging pose interaction scenarios.
Juncheng Zhu, Haotian Yang, Mubai Li et al.· Electronics· 0 citations
Experimental results on the CUFS and CUFSF datasets demonstrate that the proposed method achieves superior visual quality outperforms or ranks second in terms of the LPIPS, FID, and FSIM metrics, validating its effectiveness in high-quality cross-domain image generation.
Lei Zhang, Houpan Zhou· International Conference on...· 0 citations
A diffusion-based framework with a stage-wise multi-condition guidance mechanism that enhances both structural and textural fidelity and compares with recent image-to-image translation and diffusion-based baselines to observe competitive performance in both visual coherence and identity preservation.
Yue Que, Xuegui Cheng, Shuqian Shi et al.· Neural Networks· 0 citations
UniVVT is presented, a unified end-to-end framework that reframes VVT as semantically conditioned video generation, eliminating mask, pose, and warping modules at inference and validating implicit semantic guidance as a simple and effective alternative to fragile geometric preprocessing for end-to-end virtual try-on.
Yushe Cao, Shikun Feng, Fei Shen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.