Author

Xiang Chen

1 paper indexed here

Fetches their full publication history.

Not the right person? Other researchers publish under this name.

Jul 2026

Zero-shot cross-domain image composition via self attention injection.

Leveraging the robust generative priors of diffusion models, image composition has achieved remarkable progress. However, existing approaches continue to grapple with a persistent dilemma: the trade-off between maintaining the structural fidelity of the source object and achieving deep stylistic harmonization with the background. We attribute this limitation to two primary factors: 1) the insufficient disentanglement of geometric structure and visual appearance in current architectures, leading to conflicts during the generation process; and 2) the reliance on global statistical alignment techniques, which merely adjust tonal distributions but fail to capture complex semantic stylistic patterns. To address these challenges, we propose a novel training-free tri-branch denoising framework that effectively decouples structure from style via attention manipulation. Specifically, we propose two core mechanisms. Semantic Injection employs self attention maps to separate an object's spatial structure from its visual appearance. Style Guidance adapts advanced attention based style transfer techniques to the composition task for the first time. Comprehensive experimental results show that our method outperforms existing state-of-the-art approaches and achieves consistent improvements in structural consistency and stylistic coherence for image composition.

Xiang Chen, Qingyi Si, Bo Wang et al. · 1 citation