Skip to content

TAMF-VTON: Texture-Aware Mask-Free Virtual Try-On via High-Fidelity Image Synthesis

Jul 2026 · arXiv.org · Vol abs/2607.14807 · 0 citations · 44 references
Computer Science

TL;DR

TAMF-VTON is presented, a texture-aware, mask-free framework that enables high-fidelity image synthesis under practical unconstrained conditions and outperforms state-of-the-art methods in both quantitative metrics and perceptual quality.

Abstract

Recent diffusion-based virtual try-on (VTON) methods remain limited by their reliance on segmentation masks, insufficient preservation of fine-grained textures, and limited support for arbitrary multi-garment compositions. Consequently, existing approaches still face significant challenges in real-world e-commerce deployment. We present TAMF-VTON, a texture-aware, mask-free framework that enables high-fidelity image synthesis under practical unconstrained conditions. Our method requires no human parsing or inpainting masks at inference time and supports diverse garment styles, categories, and quantities, enabling the simultaneous transfer of multiple items while preserving body structure and intricate texture details. This is achieved through a unified generative pipeline with three key components: (1) a lightweight Mixture-of-Experts (MoE) adaptation scheme that enables efficient fine-tuning without compromising the base model's general editing capabilities; (2) a frequency-domain supervision mechanism that explicitly optimizes high-frequency spectral consistency to preserve high-fidelity textures; and (3) a robust data curation pipeline employing an adaptive inpainting strategy to simulate the inverse VTON process for high-quality training pair generation. Extensive experiments demonstrate that our approach outperforms state-of-the-art methods in both quantitative metrics and perceptual quality. Optimized for efficiency, the model achieves inference in under 15 seconds per image on an NVIDIA RTX 4090 with INT4 quantization. By combining mask-free operation, flexible multi-garment composition, faithful texture preservation, and efficient inference on consumer hardware, TAMF-VTON demonstrates a commercially viable solution for scalable deployment in real-world digital fashion scenarios. The project is available at https://www.style3d.ai/ai-photoshoot/virtual-clothing-try-on.

View source

Similar papers

#diffusion models Preprint Sep 2026

SceneHI: High-Resolution 3D-Consistent Scene Texturing with Controllable Illumination

This work introduces an exact analytical pixel-to-texel mapping that aligns diffusion trajectories across multiple viewpoints, and utilizes High-Resolution Latent Textures as a persistent canvas for gradually denoised textures, while camera views perform the denoising steps in latent pixel space.

Unknown authors · 0 citations
Book Open access Jul 2026

PixTex: Consistent 3D Texturing via Pixel-Space Multi-View Diffusion

Existing texture generation methods rely heavily on latent diffusion models, whose VAE-based spatial compression inherently limits fine-grained detail preservation and degrades pixel-level multi-view consistency. To address this limitation, we introduce PixTex, the first pixel-space multi-view diffusion framework for texture generation, which achieves substantially improved multi-view consistency. Operating directly in image space avoids latent compression, reduces inconsistencies introduced during latent-to-RGB upsampling, and preserves lossless pixel-level geometric guidance for accurate multi-view consistency. However, directly applying pixel-wise attention across multiple views is computationally prohibitive. To balance efficiency and fidelity, we adopt a coarse-to-fine consistency strategy: i) At a coarse patch level, we establish cross-view structural correspondence by employing 5D RoPE to correlate 2D patch coordinates with 3D world-space positions. ii) At the pixel level, a specialized 3D position-aware detailer further refines textural details based on patch features, ensuring fine-grained alignment unattainable by VAE-based methods. Additionally, we propose a novel consistency loss to explicitly guarantee multi-view coherence. Finally, we incorporate a pixel-space multi-view inpainting module to resolve self-occlusions and improve texture completeness. Extensive experiments demonstrate that our framework achieves state-of-the-art multi-view consistency, producing high-fidelity and seamless textures.

Yuqing Zhang, Yan-Pei Cao, Hao Xu et al. · 0 citations
Open access Aug 2026

Masking-Guided Structure and Texture Decoupling for Lightweight Blind Screen Content Image Quality Assessment

Screen content images (SCIs) exhibit complex structural heterogeneity, rendering traditional statistics-based natural scene image quality assessment (NR-IQA) metrics ineffective. Although deep learning models achieve high prediction accuracy, their prohibitive computational demands preclude deployment in latency-sensitive industrial scenarios. While existing handcrafted lightweight SCI-IQA metrics reduce computational overhead, most rely on unsegmented global feature pooling or holistic edge statistics (e.g., edge histograms or Fisher vector coding), thereby diluting locally critical text-edge distortions in vast homogeneous backgrounds. To address this limitation, we propose an ultra-lightweight, deep-learning-free NR-IQA framework centered on human visual masking. Unlike existing lightweight methods, our approach explicitly employs dual-scale Canny edge operators to partition SCIs into edge-sensitive and flat background regions. Guided by this visual prior, structural degradations and micro-compression textures are extracted region-wise using Sobel gradients and uniform local binary patterns (LBPs) and aggregated with global Commission Internationale de I’Eclairage L*a*b*(CIELAB) color statistics into a compact 60-dimensional descriptor. A grid-search-optimized Support Vector Regression (SVR) maps these features to subjective quality scores. Extensive cross-validation on the SIQAD and SCID datasets demonstrates that our metric outperforms existing handcrafted lightweight SCI metrics and traditional NSS models, while achieving accuracy competitive with representative full-reference metrics. Consuming only 79.3 ms per image on a standard CPU, it offers a practical accuracy–efficiency trade-off for resource-constrained periodic quality monitoring.

Weipeng Wu, Juan Zhang, Xiaojie Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.