Findings establish multimodal VA estimation as a soft biometric modality complementary to conventional face recognition as a soft biometric modality complementary to conventional face recognition.
Abstract
Conventional face recognition relies on static appearance cues and degrades in unconstrained settings with expression variation, occlusion, and poor lighting. We hypothesize that audiovisual expression dynamics carry identity-discriminative information complementary to static appearance, and that extracting this signal requires multimodal representations robust to the variable input quality of in-the-wild video. To learn such representations, we cast multimodal valence-arousal (VA) estimation as a pretext task and propose Quality-Aware Adaptive Fusion (QAAF), which estimates per-sample, per-modality reliability and adapts each modality's contribution through learned soft gating and a quality-dependent dropout. For the problem of VA estimation, QAAF achieves an average Concordance Correlation Coefficient (CCC) of 0.472 via late fusion ensembling on Aff-wild2, improving over a baseline ensemble under the same setting (0.415) as well as a single-backbone baseline (0.288). Furthermore, the proposed QAAF demonstrates greater resilience to unavailable modalities, with only a 7.5-34.4% relative decrease in CCC when one modality is missing. We then probe whether these VA-trained features encode identity without identity-specific training. On AFEW-VA (67 actors) and YTF (1,595 subjects), VA-trained backbone features rank first among evaluated soft biometric methods, and score-level fusion with ArcFace lowers EER on both datasets (0.022 to 0.021 on AFEW-VA, 0.106 to 0.104 on YTF), correcting 68.2% of ArcFace's false accepts on AFEW-VA. These findings establish multimodal VA estimation as a soft biometric modality complementary to conventional face recognition.
Emotion analysis is a fundamental task in computer vision, but its practical deployment remains constrained by the privacy risks inherent to conventional RGB cameras. Bio-inspired event cameras present a promising hardware-level solution because they capture asynchronous brightness changes, thereby reducing exposure of...
Jia-Qi Chen, Qin-Fu Xu, Hao Zhuang et al.· 0 citations
Dynamic Facial Expression Recognition (DFER) has recently attracted significant interest due to its vital role in enabling empathetic and human-compatible technologies. Developing models that remain robust under in-the-wild variability is a key motivation for DFER research and its practical applications. Improving mode...
Kang-Bo Ning, Shan-Shan Gao, Zhao-Qiang Xia et al.· Italian National Conference...· 0 citations
FARM-FER is proposed, which treats local and global frequency descriptors as a control signal rather than an additional classifier input, supporting a lightweight yet effective design in terms of model size and arithmetic cost for noisy-label FER.
Due to the presence of semantic ambiguity among similar expression categories and the inherent imbalance in spatio-temporal feature intensities, dynamic facial expression recognition (DFER) in the wild poses significant challenges for building trustworthy and robust systems. These factors often lead to inconsistent fea...
Feng-Qi Cui, Anyang Tong, Jinyang Huang et al.· IEEE Transactions on Informa...· 0 citations
Control comparisons and ablations indicate that the retained model has the most favorable observed cleanness–robustness trade-off among the tested epoch-matched alternatives; however, fixed-checkpoint comparisons on Occlusion-RAF-DB are not significant after Holm correction, while broader cross-domain validation remain...
Xue-Feng Zhao, Yi-Xuan Dong, Zhao-Man Zhong et al.· Italian National Conference...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.