Skip to content

AnyStyle: A Single LoRA is Sufficient for Image-Guided Style Transfer

Jul 2026 · arXiv.org · Vol abs/2607.04677 · 0 citations · 61 references
Computer Science

TL;DR

AnyStyle is proposed, a streamlined framework for image-guided style transfer that adopts a unified single-adapter paradigm for coherent style capture from the style image and incorporates training-free structural guidance from the content image, thus avoiding complex entanglement between multiple adapters and improving controllability and stability.

Abstract

Image-guided style transfer aims to apply the artistic characteristics of a style image to a content image while preserving its semantic structure and layout. Despite advances in diffusion-based methods, existing approaches often face challenges in disentangling content and style, particularly when independently optimized adapters are naively combined, causing conflicts between adapters and limiting controllability over the content-style balance in inference. We further demonstrate that training-free structural guidance directly derived from the content image through the internal attention of pre-trained model outperforms a dedicated content LoRA adapter in terms of structural fidelity and computational efficiency. Building on these observations, we propose AnyStyle, a streamlined framework for image-guided style transfer. The framework adopts a unified single-adapter paradigm for coherent style capture from the style image and incorporates training-free structural guidance from the content image, thus avoiding complex entanglement between multiple adapters and improving controllability and stability. Extensive experiments show that our method delivers competitive quantitative performance and significantly improved perceptual quality. Code is available at https://github.com/Yvan1001/AnyStyle.

View source

Similar papers

Open access Aug 2026

Optimizing Digital Art Style Transfer Using a Diffusion Model

Current generative models are prone to content distortion when guided by strong digital art styles, limiting the preservation of structural and semantic information during style transfer. Such controllable visual representation and semantic consistency are also important for intelligent imaging, computational perception, and electromagnetic information processing systems that require reliable feature preservation across heterogeneous visual domains. To address this issue, this study proposes a latent diffusion model (LDM)-based framework for optimized digital art style transfer. Strong style is quantified by the normalized norm of the CLIP-encoded style embedding in the latent space, and its distribution is employed as a continuous indicator to regulate style injection during diffusion. A Cross-Attention mechanism is introduced into the inverse denoising process to dynamically fuse content and style features while preventing structural distortion caused by excessive style guidance. In addition, a linear strength factor enables continuous adjustment of style intensity, and a pre-trained ControlNet generates regional guidance masks to spatially constrain style propagation in the latent space. Guided by style encoding, the predicted noise is iteratively refined under a fixed diffusion schedule to recover high-quality latent representations while preserving semantic consistency. Experiments on the WikiArt and COCO datasets demonstrate that the proposed method effectively alleviates content distortion and achieves a PSNR of 27.6 ± 1.2 dB and an SSIM of 0.85 ± 0.02. The proposed framework provides an effective solution for controllable image generation and offers potential value for computational imaging, intelligent visual sensing, and electromagnetic information processing applications requiring high-fidelity semantic representation.

D. Liu, L.-C. Si · 0 citations
Jul 2026

OmniStyle-INR: Universal and Multimodal Style Transfer for INRs

Style transfer remains a fundamental and highly important task across various data modalities, enabling creative manipulation conditioned by both reference images and textual descriptions. Recently, methods utilizing Gaussian Splatting have emerged as a unified representation for 2D images, video, 3D scenes, and 4D dynamics. However, representing videos and 2D images with Gaussian Splatting is structurally sub-optimal for dense continuous domains. The number of required Gaussians often approaches the total number of pixels, raising questions about the actual utility of such a representation for these specific modalities. In contrast, Implicit Neural Representations have established themselves as a much more popular and natural choice across all these data domains. Implicit Neural Representations naturally provide significant advantages, including data compression, inherent capabilities for super resolution, and seamless integration with deep generative models. To this end, we introduce OmniStyle-INR, a novel framework that leverages network-based continuous representations as a truly universal domain. Our approach successfully performs high-quality style transfer across all visual modalities, guided seamlessly by both text prompts and visual exemplars.

Rafał Kajca, Michał Miziołek, Kornel Howil et al. · 0 citations
Preprint Aug 2026

Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision

A comprehensive hierarchical taxonomy featuring over 1,000 fine-grained edit concepts is established and a dense supervision training strategy that synthesizes multiple non-interfering concepts into single image pairs is proposed that significantly enhances both training efficiency and overall model performance.

Long Cui, Xiaoqian Liu, Qi Qin et al. · 0 citations
Jul 2026

Region-aware diverse image stylization: enhancing fidelity and diversity through object-background augmentation

The region-aware diverse stylization (RDS) method is proposed, which generates multiple distinct stylized images from a single-style image without additional training and significantly outperforms state-of-the-art approaches in both fidelity and diversity.

Yang Wen, Yu-Hang Zhuang, Wuzhen Shi et al. · 0 citations
Preprint Aug 2026

Through Van Gogh's Eyes: Global Style Transfer with Diffusion Model

Artistic image synthesis aims to recreate the expressive visual identity of a target artist, yet existing methods often fail to capture an artist's global style. Conventional style transfer methods transfer the style of one or a few reference artworks to a content image in a One-to-One manner, making them effective for artwork-level stylization but limited in representing the broader stylistic distribution of an artist. Text-to-image diffusion models conditioned on artist names, such as'~ in Van Gogh style', offer greater flexibility, but they often suffer from text-induced bias and reproduce patterns from only a few iconic works. To address these limitations, we introduce Global Style Transfer (GST), an artistic image synthesis paradigm, in a Many-to-One manner, that aggregates multiple artworks from a target artist and transfers their shared global style to a single content image. For GST, we propose Global Style Guidance (GSG), which learns a residual global style offset in the intermediate feature space, or h-space, of a diffusion model under a fixed prompt. By learning artist-level style semantics purely from visual statistics, GSG mitigates text-dependent artistic bias. We further propose Content Alignment Guidance (CAG), a training-free perceptual guidance mechanism that preserves the semantic structure of the content image while allowing artist-specific geometric deformation. Experiments on WikiArt demonstrate that GST achieves superior stylistic fidelity, content preservation, and output diversity compared to existing style transfer and diffusion-based artistic synthesis methods.

J. Lee, Yujin Kim, Ghazanfar Ali et al. · 0 citations
Preprint Aug 2026

Scale-Separated Conditioning for Style-Encoder-Free Diffusion Stylization

Reference-based diffusion stylization requires separating target geometry from transferable appearance. Existing tuning-based methods often rely on aligned content-style-target triplets or auxiliary visual encoders, which increases data cost and can transfer unintended scene structure from the style reference. We propose SEFS (Style-Encoder-Free Stylization), a style-encoder-free conditioning framework for diffusion transformers. SEFS forms style tokens from stochastic low-resolution crops of single training images. This crop bottleneck preserves local appearance statistics such as palette, stroke, texture, and material, while reducing access to global layout cues. Target content is encoded by edge and segmentation cues and fused with the noisy latent through parameter-efficient trainable projections. We add style-to-denoising re-normalization for token-statistic alignment and cross-block skip fusion for spatial detail. SEFS trains on unpaired single images; the frozen diffusion VAE is used only to place image conditions in the latent space. On artistic stylization benchmarks, SEFS improves content consistency and leakage diagnostics while retaining reference-style affinity, and ablations support the crop-resolution, re-normalization, and skip-fusion choices. The code of SEFS will be made publicly available.

Jingtao Zhang, Haorui Gao, Youqin Liang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.