Skip to content

Progressive feature-space alignment for pose-controllable virtual try-on

Jul 2026 · The Visual Computer · Vol 42 · 0 citations · 44 references
Computer Science

TL;DR

A feature-space alignment framework for person-to-person multi-pose virtual try-on, which estimates appearance flow in a latent feature space and uses the learned flow to warp body-part and garment-related representations to improve realism and structural fidelity.

View source

Similar papers

Jul 2026

PIXIE: A Zero-Shot texture-invariant 6D pose estimation framework for unseen objects with assembly defects

PIXIE is a zero-shot framework that estimates the 6D pose of an object from an RGB image using only an untextured 3D model, inherently robust to lighting and texture variation, while correspondence filtering handles geometric deviations between the model and physical object.

Leon Jungemeyer, A. Magaña, Gautham Mohan et al. · 0 citations
Preprint Aug 2026

Foundational feature fusion for conditional flow matching in 6D pose estimation

Conditional flow matching has enabled a step forward in object 6D pose estimation, achieving state-of-the-art performance by progressively denoising and registering object representations to observed scenes. Existing methods require training task-specific encoders supervised on object-scene overlap and rely on trivial feature fusion strategies to resolve pose ambiguities. We present FunFlow6D, a novel flow matching-based formulation that leverages features from geometric and appearance foundation models for pose estimation, eliminating the need for task-specific encoder training. We also introduce a cross attention-based fusion mechanism that dynamically combines geometric and appearance features to provide richer conditioning for the flow matching module. Experiments on four datasets from the BOP benchmark show that FunFlow6D outperforms the previous state of the art while reducing supervision requirements and memory overhead. Extensive ablations validate the contribution of each proposed component. Project website: https://tev-fbk.github.io/FunFlow6D/.

Amir Hamza, Davide Boscaini, Fabio Poiesi · 0 citations
Preprint Aug 2026

UniVVT: A Unified End-to-End Framework for High-Fidelity Video Virtual Try-on

UniVVT is presented, a unified end-to-end framework that reframes VVT as semantically conditioned video generation, eliminating mask, pose, and warping modules at inference and validating implicit semantic guidance as a simple and effective alternative to fragile geometric preprocessing for end-to-end virtual try-on.

Yushe Cao, Shikun Feng, Fei Shen et al. · 0 citations
Aug 2026

Universal Representation for Real-World Misaligned Infrared-Visible Image Fusion.

Infrared and visible image fusion is pivotal for robust visual perception across all weather conditions and scenes. Although deep learning-based methods have made notable progress, most either assume pre-aligned inputs or rely on implicit feature-space alignment, which fails to fundamentally address the amplification of registration errors and the loss of semantic structure in the fused results. To this end, we propose a universal representation and end-to-end framework for jointly registering and fusing unaligned infrared-visible image pairs, dubbed URMIF. Each image is mapped into modality-invariant (homogeneous) and modality-specific (heterogeneous) features: the invariant "structural skeleton" encodes geometry and semantics to stabilize alignment, while the specific "texture carrier" preserves thermal saliency and visible details to enable complementary fusion. Therefore, we propose a bi-directionally coupled registration-fusion module. This module performs hierarchical deformation estimation from coarse to fine, effectively mitigating visual mismatches caused by complex parallax in real-world scenes. Within this framework, the fusion component acts as the "evaluator" of registration, providing feedback regularization to update the deformation and suppress error accumulation. Furthermore, we introduce a dominant-plane prior as a scene-level constraint, seeding stable global and patch-wise homographies and reconciling cross-modal detail conflicts, to reinforce geometric consistency and semantic reliability. We also release a large-scale dataset comprising 1,500+ unaligned infrared/visible pairs with registration ground truth, spanning diverse illumination conditions and fields of view. Based on this dataset and additional benchmarks, extensive experiments validate that our framework achieves robust alignment and high-quality fusion on misaligned inputs, markedly reducing artifacts and improving the performance of downstream tasks such as detection and segmentation. Code and benchmark are available at https://github.com/ZengxiZhang/URMIF.

Jinyuan Liu, Zengxi Zhang, Jiahao Zhang et al. · 0 citations
Open access Sep 2026

Hybrid representation and adaptive multi-feature fusion for monocular 6D pose estimation in industrial assembly

Vision-based 6D pose estimation is critical for augmented reality-assisted assembly, human–robot collaboration, and quality inspection in intelligent manufacturing. However, performance degrades severely in complex industrial scenarios due to occlusion, varying lighting, textureless surfaces, and reflective parts. This work presents a monocular 6D pose estimation approach using hybrid representations and adaptive multi-feature fusion to address these challenges. A hybrid representation learning framework is designed to jointly predict keypoint heatmaps, relational vectors, semantic edges, masks, and visibility, thereby enhancing feature robustness. A multi-feature adaptive fusion strategy optimizes the pose by combining semantic and fine-grained general features. A structure-constrained correction module refines multi-object poses using assembly consistency constraints. Experiments on a custom industrial assembly dataset and the public Mono6D dataset show that the proposed method achieves 87.46% ADD (0.1d) and 86.82% 5 cm/5∘\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$^\circ $$\end{document} accuracy, outperforming state-of-the-art methods. The custom dataset includes multiple weakly textured and reflective assembly parts under occlusion, lighting variation, and multi-viewpoint conditions. Furthermore, the complete system runs at approximately 18 FPS, with faster tracking once initialized. The approach supports reliable AR-assisted assembly and meets industrial deployment requirements. Our code and datasets are open-sourced at https://github.com/nengbinlv/HRMFPose, with the DOI: https://doi.org/https://doi.org/10.5281/zenodo.19574143.

Neng-Bin Lv, Zhang-Mao Xu, Yi Feng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.