Identity-Consistent Expression Fields (ICEF) is proposed, a framework that explicitly disentangles a static, identity-specific radiance component from a dynamic, expression-conditioned deformation component, and introduces an identity preservation regularizer that constrains the deformation network to modify only expression-relevant regions while leaving identity-specific canonical appearance untouched.
Abstract
Neural Radiance Fields (NeRF) have enabled photorealistic novel-view synthesis of 3D scenes and, in the facial domain, have been extended to reconstruct and animate 3D face models from a small number of images. However, existing few-shot dynamic NeRF methods for facial expression editing typically warp a single learned feature volume conditioned on target expression parameters, which can cause identity-specific appearance details (skin texture, fine geometric structure) to drift when the model is driven toward expressions far from those seen in the few-shot input set. We propose Identity-Consistent Expression Fields (ICEF), a framework that explicitly disentangles a static, identity-specific radiance component from a dynamic, expression-conditioned deformation component, and introduces an identity preservation regularizer that constrains the deformation network to modify only expression-relevant regions while leaving identity-specific canonical appearance untouched. ICEF further incorporates a confidence-weighted conditional feature warping step that down-weights unreliable warps for target expressions that are far, in parameter space, from the observed few-shot inputs, mitigating artifacts observed in prior few-shot dynamic NeRF methods when extrapolating to novel expressions. We relate ICEF to prior few-shot dynamic NeRF, static 3D-aware face generation, and disentangled face-editing radiance field methods, and describe an evaluation protocol measuring both novel-expression rendering quality and, specifically, identity-consistency metrics across a range of expression-parameter extrapolation distances.
—The reconstruction of photorealistic 4D faces from monocular video holds significant importance for applications such as virtual reality and digital human-robot interaction. However, achieving reconstructions with structurally reasonable facial geometry and natural expressions remains challenging. Existing approaches typically combine dynamic neural radiance fields (NeRFs) with low-dimensional parametric models, such as 3D morphable model (3DMM), for facial reconstruction and expression control. Nevertheless, 3DMM-based expression features often lack the ability to preserve high-frequency details and subtle individual variations. To address this issue, we propose CEE-NeRF, a novel framework that integrates geometry-aware 3DMM features with high-level semantic features through contrastive learning for complementary feature fusion. This design enhances the representation of fine-grained facial details while mitigating expression artifacts. In addition, unlike previous studies, we propose reformulating monocular 4D face reconstruction as a multimodal learning problem and replacing the conventional multilayer perceptron used in NeRF with a mixture-of-experts Transformer, which enables modality-specific expert processing for accurate density and color estimation. This multimodal framework is also highly extensible and can be easily adapted to additional modalities. Extensive experiments demonstrate that CEE-NeRF achieves high-fidelity reconstructions from challenging monocular videos, consistently outperforming state-of-the-art approaches in rendering quality.
Hao Yu, Hui Yu, Rachael E. Jack et al.· IEEE Transactions on Computa...· 0 citations
These results establish frozen DINOv3 as a strong zero-shot representation for region-level facial correspondence and identify intermediate self-supervised features as the most useful layer for dense face analysis.
Izaldein Al-Zyoud, Abdulmotaleb El Saddik· arXiv.org· 0 citations
Dynamic 3D Gaussian Splatting (3DGS) enables efficient facial-avatar rendering but may suffer from cumulative geometric drift and degradation of identity-specific details when applied to long monocular sequences. To address this problem, we propose IA-3DGS, an identity-anchored dynamic Gaussian framework for long-term facial animation. The proposed framework contains three main components. First, a 3D morphable model-guided semantic initialization strategy associates Gaussian primitives with anatomically meaningful facial regions. Second, a region-weighted Jacobian rigidity regularizer suppresses non-physical shear and anisotropic stretching in quasi-rigid facial regions while preserving sufficient flexibility in expression-sensitive regions. Third, a cyclic memory correction mechanism periodically aligns the current identity representation with a frozen reference representation and applies exponential moving average smoothing to reduce correction-induced temporal discontinuities. Experiments were conducted on the NeRSemble and NHA datasets and on a self-collected Custom-5K dataset containing continuous monocular facial sequences longer than 5000 frames. IA-3DGS achieved a PSNR of 31.85 dB, an SSIM of 0.942, and an LPIPS of 0.038 on standard-length sequences. At frame 5000, the method obtained an L-LPIPS of 0.046 and a landmark mean error of 1.3 mm, compared with 0.245 and 6.8 mm, respectively, for the purely implicit dynamic 3DGS baseline. The system rendered 1920 × 1080 images at approximately 48 frames per second on a single NVIDIA RTX 4090. These results indicate that the proposed semantic, geometric, and temporal constraints improve long-sequence identity consistency under the evaluated subject-specific monocular setting.
Ze-Wei Zhao, L. Yin, Shuaijie Wang et al.· Italian National Conference...· 0 citations
3D Gaussian Splatting (3DGS) delivers real-time and high-fidelity rendering but remains challenged by unconstrained in-the-wild scenes, where drastic appearance variations and transient objects violate multi-view consistency. Existing methods are fundamentally limited by independent and discrete embeddings that struggle to capture continuous environmental changes or model spatially-varying local illumination. To address these limitations, we propose \textbf{WilLaGS}, a unified framework for robust 3D scene reconstruction and generative appearance synthesis under unconstrained settings. Specifically, we introduce a generative appearance model where a $\beta$-VAE learns a structured and continuous manifold of global appearance. Conditioned on the latent code, we construct a 3D neural appearance field that generates dynamic Tri-Plane features to encode spatially-varying local illumination effects. Furthermore, to suppress transient artifacts, we present a self-supervised perceptual masking mechanism that leverages a Teacher-Student (EMA) architecture to derive a stable scene consensus, robustly identifying inconsistent regions via perceptual discrepancies. Extensive experiments on multiple datasets demonstrate that \textbf{WilLaGS} achieves state-of-the-art performance in reconstruction quality and novel view appearance synthesis, while maintaining real-time rendering efficiency.
Yu Bai, Qian-Qiu Tan, Li-Long Chen et al.· 0 citations
Creating re-topologized 3D facial meshes is essential for high-quality facial animation but remains labor-intensive and time-consuming. This dissertation explores more efficient approaches for capturing production-ready facial meshes through: (1) the development of VarIS, a custom light sphere for capturing high-resolution stereo geometry and reflectance maps; (2) analysis of camera parameters affecting automatic 2D and 3D landmarking; (3) synthetic-data methods for training neural face regression; and (4) techniques for improving neural multi-view face-shape regression. While VarIS enables photorealistic face capture, its operational and processing costs motivate a more scalable approach. A deep learning framework is therefore proposed to directly predict re-topologized facial meshes from synthetic multiview images generated with Visage Craft, an in-house physically based rendering system using an Appearance 3D Morphable Model (A3DMM). The system produces standardized meshes ready for rigging and animation with minimal human supervision. Results show that incorporating accurate camera intrinsics and extrinsics improves landmark accuracy and geometric consistency, while 3D landmark regularization further improves reconstruction quality.
Xiang Li· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.