HTNet-SDGCN: hierarchical transformer with structural-dynamic dual graph convolutional network for micro-expression recognition
Abstract
Micro-expression recognition (MER) is a fundamental yet challenging task in affective computing due to the subtle, transient, and localized nature of spontaneous facial muscle movements. Although Vision Transformers effectively extract MER features, spatial regions are often processed in isolation, limiting the ability to capture long-range relational dependencies across facial zones. To address this problem, this study proposes a so-called HTNet-SDGCN that augments a hierarchical Transformer backbone with a Structural-Dynamic Dual Graph Convolutional Network (SD-GCN). SD-GCN models inter-region dependencies via two complementary graphs, i.e., a shared structural graph that learns stable facial topology across all training samples, and a dynamic graph that adapts to the motion pattern of each individual input. These graphs are adaptively fused to generate a comprehensive relational representation. Furthermore, Landmark-Guided Gaussian Weighting is introduced to sharpen optical-flow signals around semantic facial landmarks. Extensive experiments on the CASME II, SMIC, and SAMM datasets under the leave-one-subject-out (LOSO) protocol demonstrate that the proposed framework achieves robust performance, attaining a UF1 of 0.8813 and a UAR of 0.8661 on the full composite dataset, and reaching a peak UF1 of 0.9564 on CASME II.