Skip to content
Open access

StyleSmith: A Signature Style Transfer Framework With Parametric and Temporal Attention

2026 · IEEE Access · Vol 14, pp. 148303-148336 · 0 citations · 49 references

Abstract

Artistic style transfer, which renders content images in the visual style of a reference artwork while preserving semantic structure, is a fundamental problem in computer vision and AIGC, with broad applications in digital art, commercial design, and creative media. Despite progress driven by diffusion-based approaches, signature style transfer, defined as the high-fidelity transfer of highly recognizable artistic traits such as exclusive geometric structures, personalized color schemes, brushstroke textures, and compositional logic, remains underexplored. Existing methods fall into two paradigms with critical limitations: training-free methods (e.g., StyleID, InstantStyle) enable efficient single-pass transfer but capture only shallow color and global texture features, failing to model fine-grained signature traits; tuning-based methods (e.g., InST) achieve stronger customization via text inversion but suffer from severe overfitting in one-shot fine-tuning, high computational cost, and content degradation. Four bottlenecks remain unresolved: insufficient fine-grained style capture, limited modeling from single reference images, an inherent style-content trade-off, and one-shot overfitting. To address these gaps, we propose StyleSmith, a diffusion framework for high-fidelity signature style transfer. First, a parametric style modulation network (PSM-Net) driven style-aware fine-tuning mechanism performs targeted optimization solely on the UNet decoder attention weights, identified as the core module for style learning, enabling accurate fine-grained style encoding from a single reference image, reducing tuned parameters by over 60%, and mitigating one-shot overfitting. Second, a temporal content preservation approach via attention anchoring (TCP-AA) injects content structural priors in early denoising stages and relaxes constraints in later stages, achieving a dynamic style-content balance analogous to an exploration-exploitation mechanism. Extensive experiments on WikiArt and ArtBench against 8 state-of-the-art methods show StyleSmith achieves the lowest style loss (0.7641, statistically significant, paired t-test $p \leq 0.033$ ) and FID (12.37), ranks second in content fidelity (LPIPS = 0.5191), and reduces fine-tuning time by nearly 50% versus comparable tuning-based methods. A user study with 25 evaluators further confirms StyleSmith’s perceptual superiority in style fidelity, content preservation, and overall quality. The framework also supports local style transfer, texture transfer, and style-guided text-to-image generation, demonstrating strong academic innovation and industrial value.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.