Optimizing Digital Art Style Transfer Using a Diffusion Model
Abstract
Current generative models are prone to content distortion when guided by strong digital art styles, limiting the preservation of structural and semantic information during style transfer. Such controllable visual representation and semantic consistency are also important for intelligent imaging, computational perception, and electromagnetic information processing systems that require reliable feature preservation across heterogeneous visual domains. To address this issue, this study proposes a latent diffusion model (LDM)-based framework for optimized digital art style transfer. Strong style is quantified by the normalized norm of the CLIP-encoded style embedding in the latent space, and its distribution is employed as a continuous indicator to regulate style injection during diffusion. A Cross-Attention mechanism is introduced into the inverse denoising process to dynamically fuse content and style features while preventing structural distortion caused by excessive style guidance. In addition, a linear strength factor enables continuous adjustment of style intensity, and a pre-trained ControlNet generates regional guidance masks to spatially constrain style propagation in the latent space. Guided by style encoding, the predicted noise is iteratively refined under a fixed diffusion schedule to recover high-quality latent representations while preserving semantic consistency. Experiments on the WikiArt and COCO datasets demonstrate that the proposed method effectively alleviates content distortion and achieves a PSNR of 27.6 ± 1.2 dB and an SSIM of 0.85 ± 0.02. The proposed framework provides an effective solution for controllable image generation and offers potential value for computational imaging, intelligent visual sensing, and electromagnetic information processing applications requiring high-fidelity semantic representation.