Skip to content

Region-aware diverse image stylization: enhancing fidelity and diversity through object-background augmentation

Jul 2026 · The Visual Computer · Vol 42 · 0 citations · 46 references

TL;DR

The region-aware diverse stylization (RDS) method is proposed, which generates multiple distinct stylized images from a single-style image without additional training and significantly outperforms state-of-the-art approaches in both fidelity and diversity.

View source

Similar papers

Conference 2026

Structure-Aware and Frequency-Guided Diffusion Framework for Multimodal Fashion Image Editing

Fashion image editing aims to modify target garment attributes under textual and reference guidance while preserving non-target contents. However, existing methods often suffer from inaccurate garment localization, insufficient preservation of high-frequency textures, and unnatural transitions near edited boundaries. To address these issues, we propose a structure-aware and frequency-guided framework for multimodal fashion image editing. Specifically, we design a Structure-Guided Textual Mask Network to predict geometry-aware editing regions by leveraging refined textual structural cues and human-centric priors, where a structural prior reweighting mechanism is introduced to improve localization accuracy. We further develop an adaptive frequency-domain texture enhancement module to inject high-frequency fabric details from a reference image during late denoising, and employ a boundary-band soft fusion strategy to ensure smooth visual transitions. In addition, we construct a new dataset, DFEdit, for fine-grained multimodal fashion image editing. Experimental results show that the proposed method achieves competitive performance in terms of editing fidelity, texture consistency, and visual quality, showing its effectiveness for intelligent fashion image editing applications.

Xin Chen · 0 citations
Jul 2026

InstructMixup: Instruction-Guided Salient Patch Editing for Robust Data Augmentation

In image and video technologies, data augmentation is widely used to improve the generalization of deep visual models, and mixup-based strategies that interpolate between samples have become the dominant approach. However, computing informative mixing regions adds substantial overhead, and blending content across different images frequently disrupts the semantic integrity of the resulting sample. We propose \our{}, a data augmentation method that constructs challenging yet label-consistent training samples entirely within a single visual sample. \our{} first extracts multi-scale salient patches from the sample using a lightweight saliency detector, refines each patch with an instruction-guided generative model, and blends the edited patch back into the non-salient regions of the same sample; because the generative edits are computed once and cached offline, this step adds negligible training cost. To further diversify the learned representation, \our{} injects self-similar fractal structure into the same salient regions at an adaptive ratio, so each training sample carries both fractal and non-fractal structure. We derive a second-order approximation of the resulting vicinal risk, showing that the method simultaneously enforces invariance to the generative edit and suppresses curvature along the perturbed salient directions, and we verify both predictions empirically. We evaluate on small to large backbones for instance Convolutional Neural Networks (CNNs), Vision Transformers (ViTs) and Vision-Language Foundational Models (VLMs) across seven benchmarks covering coarse- and fine-grained classification, robustness to corruption and occlusion, calibration, and transfer and self-supervised learning, InstructMixup outperforms nine competing augmentation methods, surpassing the strongest baseline across all benchmarks.

Khawar Islam, Arif Mahmood, Xin Jin et al. · 1 citation
Preprint Aug 2026

Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision

A comprehensive hierarchical taxonomy featuring over 1,000 fine-grained edit concepts is established and a dense supervision training strategy that synthesizes multiple non-interfering concepts into single image pairs is proposed that significantly enhances both training efficiency and overall model performance.

Long Cui, Xiaoqian Liu, Qi Qin et al. · 0 citations
Open access Aug 2026

Using Image Style Transfer Technology Based on U-Net Architecture to Improve the Information Expression Effect of Intangible Cultural Heritage Brand Visual Posters

Traditional poster design often struggles to efficiently and accurately preserve the complex styles, textures, and cultural connotations of intangible cultural heritage (ICH), resulting in limited visual information expression. To address this challenge, this paper proposes an adaptive feature fusion style transfer model based on the U-Net architecture, which provides a high-fidelity visual information processing framework with potential methodological implications for intelligent image transmission and multi-scale feature representation in advanced engineering systems. The proposed model exploits the symmetric encoder–decoder structure and skip connections of U-Net to achieve cross-scale extraction and transfer of ICH style features while preserving fine content details. An Adaptive Instance Normalization (AdaIN) module is employed to decouple and fuse style and content features in the latent space, enabling efficient style injection without sacrificing structural consistency. Skip connections further facilitate the integration of high-resolution semantic information with stylized representations, producing visual posters that simultaneously exhibit cultural authenticity and clear information delivery. Experimental results demonstrate superior performance, achieving a Style Consistency Index (SCI) of 0.28, an Information Clarity Metric (ICM) of 0.95, and a Brand Recognition Perception Score (BRPS) of 4.50, outperforming existing comparison methods. The proposed approach effectively improves both artistic quality and information expression efficiency and provides a reliable feature fusion strategy for intelligent visual communication and digital information processing applications.

Y. Yuan, Y. Liu · 0 citations
Preprint Aug 2026

Staying True to the Origin: Continuous Image Stylization with Smooth Transitions

Recent advances in generative models have achieved remarkable performance in text- and image-conditioned editing. However, preserving the content of a given image while referencing style patterns from another remains challenging, often leading to uncontrollable stylization results. In this paper, we approach image stylization from the perspective of continuous control, aiming to enable modern Diffusion Transformer (DiT)-based multi-reference editing models to (1) faithfully preserve the semantic structure of the content image, (2) render strong stylization effects, and (3) smoothly transition between the two. To this end, we propose a simple yet effective two-stage training strategy along with a style-strength-aware spline formulation. Specifically, in the first stage, the model is trained to produce strongly stylized outputs while preserving the content semantics as much as possible. In the second stage, with the base model frozen, we learn a set of anchor projectors that map various stylization strengths into the model parameter space. During inference, by performing style-strength-aware spline interpolation in a low-rank space, our method enables continuous control over stylization strength, even though the model is trained with only a few discrete strength levels. Extensive experiments demonstrate that our method supports precise and continuous manipulation of stylization strength while generating high-fidelity results with modern DiT models. Project page: https://reychiaro.github.io/StyleController.

Rui Xu, Han Zhang, Songhua Liu · 0 citations
Jul 2026

InnoText: A Unified Model for Visual Text Generation and Editing

This work proposes InnoText, a unified DiT-based framework capable of performing both text generation and editing within a single model, and introduces a Font Size-Aware Modulation module to enhance representations across font scales, a Small-Character Aware Augmentation strategy to improve fine-grained fidelity, and a Task-Specific Region Weighted Loss for adaptive optimization.

Hao-Wei Liu, Runze He, Jian Lu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.