Skip to content
Book Open access

VFAvatar: Feed-Forward 3D Avatar Reconstruction from Casual Image Collections

Jul 2026 · International Conference on Computer Graphics and Interactive Techniques · 0 citations · 45 references
Computer Science

Abstract

The growing demand for personalized 3D avatars calls for efficient reconstruction methods from casual photos. This task remains challenging due to unconstrained viewpoints, partial body visibility, and temporal variations across input images. While some previous methods circumvent these difficulties by adopting generative approaches like score distillation, they struggle to preserve authentic appearance details from source images. To address these limitations, we introduce Visual-Fusion-Avatar (VFAvatar), a novel feed-forward framework that reconstructs 3D avatars by fusing visual cues in just a few seconds. VFAvatar couples a pose-free reconstruction foundation model with a pretrained human generation prior in a mutually reinforcing manner. And we propose a visibility-aware, view-attentive residual aggregation mechanism that routes and fuses per-view updates, allowing partial observations from different images to be assembled into a single coherent avatar. Experiments demonstrate that VFAvatar significantly outperforms state-of-the-art methods in both reconstruction fidelity and efficiency, while enabling shape and pose manipulation. Code is available on: https://github.com/huangshuo200823/VFAvatar.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.