Skip to content

Causal Supervision of Attention for Affective Behaviour Analysis

Jul 2026 · arXiv.org · Vol abs/2607.12091 · 0 citations · 48 references
Computer Science

TL;DR

An attention pooling framework that combines causal supervision with cross-covariance regularization of attention components is proposed, encouraging subject-invariant attention and non-redundant representations that improve generalization.

Abstract

The \textit{11th Affective Behaviour Analysis in-the-wild Competition} includes the Multi-Task Learning Challenge, where participants develop a unified framework for Valence-Arousal Estimation, Expression Recognition, and Action Unit Detection. The challenge lies in learning emotion-related representations that generalize across subjects while remaining robust to spurious factors such as identity, illumination, pose, and demographic variation. To aggregate features extracted by a pre-trained backbone into a compact representation for prediction, attention mechanisms selectively weight the most informative facial regions. However, these attention weights can still capture dataset-specific correlations rather than genuine affective cues. To address this limitation, we propose an attention pooling framework that combines causal supervision with cross-covariance regularization of attention components, encouraging subject-invariant attention and non-redundant representations that improve generalization. Our method achieves $CCC_{VA}=0.5123$ for VA estimation on the official validation set, together with $F_{EX}=0.3116$ and $F_{AU}=0.3974$ for expression recognition and action unit detection, respectively, resulting in an overall $P$ score (the sum of the individual task metrics) of $1.2214$.

View source

Similar papers

Jul 2026

AffectFuse: Cross-Task Feature Fusion with Temporal Modeling for Multi-Task Affective Behavior Analysis

This work presents a system for the Multi-Task Learning (MTL) track of the 11th Affective Behavior Analysis in-the-wild (ABAW) competition on s-Aff-Wild2, the static selected-frame version of Aff-Wild2, showing that post-encoder adaptation and task-wise modeling choices provide a strong MTL pipeline without training a...

Dipit Saha, Mohammad Raihan Rashid, Shahruz Mannan et al. · 0 citations
Open access Aug 2026

Emotion recognition from body movement through interpretable motion-aware sequential modeling

Emotion recognition from bodily movement remains a challenging problem, particularly when only pose-based motion sequences are available and emotionally informative content is not uniformly distributed across time. In this work, we propose a Window Transformer architecture grounded in the Multiple Instance Learning (MI...

Sergio Esteban-Romero, Iván Martín-Fernández, R. San-Segundo et al. · 0 citations
Preprint Aug 2026

Mind the Student: Behavioral and Contextual Cues for Automated Engagement Prediction in Online Learning

The prediction of student engagement from the online tutoring videos is difficult because engagement is a multidimensional construct comprising distinct behavioral, emotional, and cognitive states. A reliable prediction requires bringing together different types of behavioral signals as well as expressive cues. Through...

Alperen Kantarcı, Visvanathan Ramesh, Gemma Roig · 0 citations
Preprint Aug 2026

LG-GER: Language-Guided Group Emotion Recognition via Multimodal Evidence Distillation

Inferring the collective emotional state of a group of people from a single image, a task known as group emotion recognition (GER), requires integrating spatially distributed cues such as faces, poses, interactions, and scene context. Current methods rely on detector-driven multi-stream pipelines. These are trained wit...

Ahmed-Shehab Khan, Zhiyuan Li, Yan Tong · 0 citations
Open access 2026

MM-PSYCHE: Multimodal Multitask Psychological Characteristic Estimation Through Cross-Domain Semi-Supervised Learning

Psychological characteristic estimation from multimodal in-the-wild behavior is usually studied using separate corpora, each annotated for a single target task. Such annotation fragmentation limits cross-task learning and cross-domain generalization across affective, dispositional, and interactional phenomena. To addre...

E. Ryumina, A. Axyonov, D. Koryakovskaya et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.