This work casts partially-labeled multi-task learning as marginalization over a shared affect latent: one variational bottleneck mediates all three task decoders, so a frame annotated for one task shapes the representation the others use, and the masked objective reappears as the reconstruction term of an evidence lower bound.
Abstract
Facial affect in the wild is naturally multi-task: valence-arousal, discrete expressions, and facial action units describe the same face. Yet real corpora annotate these tasks only partially and unevenly, so most systems mask the missing labels or impute pseudo-labels and forgo the cross-task signal. We instead cast partially-labeled multi-task learning as marginalization over a shared affect latent: one variational bottleneck mediates all three task decoders, so a frame annotated for one task shapes the representation the others use, and the masked objective reappears as the reconstruction term of an evidence lower bound. On s-Aff-Wild2, where only 37% of frames carry all three labels, the classes are severely imbalanced, and pretraining on the source data is disallowed, we isolate where this coupling acts. On a single backbone it lifts expression macro-F1 from 0.403 for a dedicated specialist to 0.446, which the masked-loss model does not reach; a second, near-peer backbone with decorrelated errors then breaks an action-unit ceiling that external action-unit data could not, while valence-arousal stays within noise. Every gain is disciplined by a matched-control negative; together these controls indicate that the rare-class failure is representational, not a matter of loss shaping. As each task's source is chosen on the evaluation split, we report the assembled result, a combined multi-task score of 1.679 on validation, as an in-sample endpoint and rest our conclusions on the controlled comparisons; a small, regime-dependent transfer of the expression advantage to AffectNet and RAF-DB is presented as exploratory rather than conclusive.
This work presents a system for the Multi-Task Learning (MTL) track of the 11th Affective Behavior Analysis in-the-wild (ABAW) competition on s-Aff-Wild2, the static selected-frame version of Aff-Wild2, showing that post-encoder adaptation and task-wise modeling choices provide a strong MTL pipeline without training a...
Dipit Saha, Mohammad Raihan Rashid, Shahruz Mannan et al.· arXiv.org· 0 citations
Control comparisons and ablations indicate that the retained model has the most favorable observed cleanness–robustness trade-off among the tested epoch-matched alternatives; however, fixed-checkpoint comparisons on Occlusion-RAF-DB are not significant after Holm correction, while broader cross-domain validation remain...
Xue-Feng Zhao, Yi-Xuan Dong, Zhao-Man Zhong et al.· Italian National Conference...· 0 citations
It is suggested that, for a lightweight 18-layer network, residual connections contribute marginally to overall accuracy but help stabilize training and improve feature learning for underrepresented classes.
The FER20E dataset provides a comprehensive benchmark for advancing emotion recognition in unconstrained and real-world scenarios, and a data annotation tool (DL-DAT) that follows a semi-automated, human-in-the-loop pipeline to enable scalable and reliable annotation.
Facial expression recognition (FER) is increasingly required in classroom affect analysis, lightweight human-computer interaction, and domain-specific behavioral monitoring, where only limited labeled facial images are available. Reliable FER in such low-resource settings is critical because unstable predictions can di...
Yang-Xuan Xie· Applied and Computational En...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.