Jul 2026· 2026 11th International Conference on Applying New Technology in Green Buildings (ATiGB)· pp. 1060-1068· 0 citations· 53 references
Abstract
Image-based virtual try-on (VTON) has advanced rapidly with the emergence of high-resolution generative adversarial networks and diffusion-based synthesis models. However, many remaining failures, including garment misalignment, boundary artifacts, unrealistic deformation, and identity or body-shape distortion, are not caused only by generator limitations but also by the quality, structure, and fusion of upstream input representations. This paper presents an input-centric survey of image-based VTON systems. Unlike prior reviews that mainly organize the field by generative architecture, this work analyzes how geometry, semantic region control, garment conditioning, and multi-modal fusion shape the final try-on output. We review representative VTON methods, datasets, and evaluation practices, and group them according to the role of pose, body representation, parsing masks, garment appearance, and conditioning signals. We further discuss how representation errors propagate through warping, synthesis, and diffusion-conditioning stages. The survey highlights three main findings: (i) representation quality places an upper bound on synthesis realism, (ii) mask and semantic-region quality remain major bottlenecks even in recent diffusion-based approaches, and (iii) garment material representation is still weakly modeled in existing pipelines. Finally, we identify open research directions toward uncertainty-aware masks, material-informed garment embeddings, standardized evaluation protocols, and robust input fusion for real-world VTON deployment.
Text-to-image generation is an increasingly fast-paced field of generative artificial intelligence, consisting of synthesizing images of high quality and semantic consistency based on natural language descriptions. In this paper, we give an extensive overview of the approach to text-to-image generation using deep learning, including the most common core model families, architecture designs, training approaches, and evaluation systems. We discuss the paradigms of the generative adversarial networks (GANs), variational autoencoders (VAEs), transformer-based designs, and diffusion models, with the last one representing the state of the art in image generation models. The review also discusses key aspects of pipelines such as text encoding, cross-modal alignment, mechanisms of attention, and decoding images. Popular datasets, methods, and metrics of evaluation, including Fréchet Inception Distance (FID) and CLIP-based similarity, are discussed. The application domains that involve creative content creation, medical imaging, education and industrial design are critically discussed. Despite significant advances, various issues still exist, such as low stability in training, excessive computational complexity, amplification of bias, generated images, and text–image alignment errors. Moral and social issues, such as misinformation, intellectual property, and equity, are critically examined. Lastly, we present future research directions to more controllable, more efficient and more interpretable text-to-image systems, focusing on multimodal foundation models and human–AI collaborative design.
A diffusion-based framework with a stage-wise multi-condition guidance mechanism that enhances both structural and textural fidelity and compares with recent image-to-image translation and diffusion-based baselines to observe competitive performance in both visual coherence and identity preservation.
Yue Que, Xuegui Cheng, Shuqian Shi et al.· Neural Networks· 0 citations
GenRec is introduced, a multi-view flow matching model that builds the reconstruction--generation split directly into its architecture, supervision, and gradient flow, and attains the best reconstruction fidelity in observed regions while also surpassing purely generative baselines on perceptual quality in unobserved ones.
Ata Çelen, Jaewoo Jung, Federico Tombari et al.· 0 citations
TAMF-VTON is presented, a texture-aware, mask-free framework that enables high-fidelity image synthesis under practical unconstrained conditions and outperforms state-of-the-art methods in both quantitative metrics and perceptual quality.
Jie Wang, Qian He, Gaofeng He et al.· arXiv.org· 0 citations
PoseAdapter, a lightweight framework for high-fidelity 2.5D controllable image generation, and a Context-Aware Dual-Stream Representation, to resolve the generative trade-off between strict instance isolation and global coherence.
Yufeng Chi, Hui-Min Ma, Fan Gao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.