Skip to content

Quality-Aware Multimodal Fusion Reveals Implicit Identity in Valence-Arousal Features

Jul 2026 · arXiv.org · Vol abs/2607.21347 · 0 citations · 50 references
Computer Science

TL;DR

Findings establish multimodal VA estimation as a soft biometric modality complementary to conventional face recognition as a soft biometric modality complementary to conventional face recognition.

Abstract

Conventional face recognition relies on static appearance cues and degrades in unconstrained settings with expression variation, occlusion, and poor lighting. We hypothesize that audiovisual expression dynamics carry identity-discriminative information complementary to static appearance, and that extracting this signal requires multimodal representations robust to the variable input quality of in-the-wild video. To learn such representations, we cast multimodal valence-arousal (VA) estimation as a pretext task and propose Quality-Aware Adaptive Fusion (QAAF), which estimates per-sample, per-modality reliability and adapts each modality's contribution through learned soft gating and a quality-dependent dropout. For the problem of VA estimation, QAAF achieves an average Concordance Correlation Coefficient (CCC) of 0.472 via late fusion ensembling on Aff-wild2, improving over a baseline ensemble under the same setting (0.415) as well as a single-backbone baseline (0.288). Furthermore, the proposed QAAF demonstrates greater resilience to unavailable modalities, with only a 7.5-34.4% relative decrease in CCC when one modality is missing. We then probe whether these VA-trained features encode identity without identity-specific training. On AFEW-VA (67 actors) and YTF (1,595 subjects), VA-trained backbone features rank first among evaluated soft biometric methods, and score-level fusion with ArcFace lowers EER on both datasets (0.022 to 0.021 on AFEW-VA, 0.106 to 0.104 on YTF), correcting 68.2% of ArcFace's false accepts on AFEW-VA. These findings establish multimodal VA estimation as a soft biometric modality complementary to conventional face recognition.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Emo-DVS: A Multimodal Benchmark for Privacy-Aware Emotion Recognition with Event Cameras

Emotion analysis is a fundamental task in computer vision, but its practical deployment remains constrained by the privacy risks inherent to conventional RGB cameras. Bio-inspired event cameras present a promising hardware-level solution because they capture asynchronous brightness changes, thereby reducing exposure of...

Jia-Qi Chen, Qin-Fu Xu, Hao Zhuang et al. · 0 citations
Open access Aug 2026

Parameter-Efficient Audio-Visual Dynamic Facial Expression Recognition with Mamba Fusion Adapters and Frame-Level Feature Arrangement

Dynamic Facial Expression Recognition (DFER) has recently attracted significant interest due to its vital role in enabling empathetic and human-compatible technologies. Developing models that remain robust under in-the-wild variability is a key motivation for DFER research and its practical applications. Improving mode...

Kang-Bo Ning, Shan-Shan Gao, Zhao-Qiang Xia et al. · 0 citations
Open access Aug 2026

Frequency-Guided Expert Modulation for Noisy-Label Facial Expression Recognition

FARM-FER is proposed, which treats local and global frequency descriptors as a control signal rather than an additional classifier input, supporting a lightweight yet effective design in terms of model size and arithmetic cost for noisy-label FER.

Miaomiao Zhang, Meng Lou, Linwei Chen · 0 citations
2026

Toward Trustworthy Dynamic Facial Expression Recognition via Information Bottleneck Modeling

Due to the presence of semantic ambiguity among similar expression categories and the inherent imbalance in spatio-temporal feature intensities, dynamic facial expression recognition (DFER) in the wild poses significant challenges for building trustworthy and robust systems. These factors often lead to inconsistent fea...

Feng-Qi Cui, Anyang Tong, Jinyang Huang et al. · 0 citations
Open access Aug 2026

Compact Occlusion-Robust Facial Expression Recognition via Clean-Anchored Hard Occlusion Fine-Tuning

Control comparisons and ablations indicate that the retained model has the most favorable observed cleanness–robustness trade-off among the tested epoch-matched alternatives; however, fixed-checkpoint comparisons on Occlusion-RAF-DB are not significant after Holm correction, while broader cross-domain validation remain...

Xue-Feng Zhao, Yi-Xuan Dong, Zhao-Man Zhong et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.