Skip to content
Open access

HiPA-Gen: Hierarchical Interaction Perception and Alignment for High-Fidelity Pose-Guided Human Generation

Aug 2026 · Electronics · 0 citations · 15 references

Abstract

Pose-guided human image generation aims to synthesize an image of a target person based on a reference image and a target pose. Although diffusion-based methods have recently achieved significant progress in pose alignment and visual realism, generating high-fidelity images in regions involving complex pose interactions remains a challenge. Typical issues include structural confusion of limbs and texture distortion in occluded areas. To address this challenge, this paper proposes the HiPA-Gen framework, which is designed to automatically perceive complex interaction regions and generate high-fidelity content within them. Unlike existing methods that mainly rely on flat skeleton conditions or implicit attention responses, HiPA-Gen explicitly models hierarchical spatial relations and local appearance priors in complex interaction regions. Specifically, we design a Dual-Agent Hierarchical Reasoning Module (DHR) where a Prompt-Guided Reasoning Agent identifies interaction-related body parts and a Hierarchical Rendering Agent converts the inferred relations into a Hierarchical Correspondence Pose (HCP) map. The HCP map provides explicit front–back structural cues for limb overlap, hand occlusion, and body self-occlusion. To further reduce local texture ambiguity, we introduce a Part Prior Detail Alignment Module (PDA), which extracts Regional Reference Assets (RRAs) from the source image under HCP-guided part priors. These regional references are then injected into the Pose-Conditioned Interaction-Aware Detail Synthesis Network (PDS) through a Local Enhancement Branch (LEB), enabling more accurate local feature fusion during diffusion-based synthesis. Experiments on DeepFashion and Market-1501 show consistent numerical improvements under the reported evaluation protocols in structural similarity, perceptual quality, and distributional fidelity. Qualitative results further show that the proposed framework produces clearer limb boundaries, more reliable spatial ordering, and more consistent clothing textures in challenging pose interaction scenarios.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.