Transformer Based Modeling of Spatio Temporal Facial Dynamics for Micro-Expression Recognition
Abstract
Micro-temporal facial expressions consist of brief and subtle muscle activations that convey genuine emotional states. Accurately modeling these transient dynamics is challenging due to their low intensity, short duration, and limited spatial variation, which often hinder the performance of conventional convolutional and recurrent architectures. To address these limitations, this paper introduces a Spatio-Temporal Self-Attention Network (STTN) that leverages transformer-based self-attention to effectively capture finegrained dependencies across both spatial regions and short temporal intervals. The proposed framework focuses on learning discriminative representations of micro-temporal facial movements by emphasizing relevant facial regions and their temporal evolution. Extensive experiments are conducted on high-frame-rate micro-expression benchmarks, including SAMM, CASME II, and CAS(ME)2 datasets. The results demonstrate that the proposed model achieves superior performance compared to existing state-of-the-art approaches, highlighting the effectiveness of self-attention mechanisms in modeling subtle and rapid facial dynamics.