Skip to content
Open access

Emotion recognition from body movement through interpretable motion-aware sequential modeling

Aug 2026 · Frontiers in Artificial Intelligence · Vol 9 · 0 citations · 58 references
Medicine

Abstract

Emotion recognition from bodily movement remains a challenging problem, particularly when only pose-based motion sequences are available and emotionally informative content is not uniformly distributed across time. In this work, we propose a Window Transformer architecture grounded in the Multiple Instance Learning (MIL) paradigm to address this challenge. Rather than processing the full sequence as a single temporal stream, the model decomposes it into overlapping windows and learns to assign greater relevance to those segments containing stronger emotional content. This formulation provides a more interpretable framework, since the learned relevance scores reveal which temporal regions drive the final prediction, while also yielding richer, context-aware representations of each segment. We evaluate the proposed approach on two publicly available datasets, MEED and DIEM-A, and compare it against a standalone Transformer baseline under different batch size and window configuration settings. The Window Transformer consistently outperforms the baseline and exhibits a more stable behavior across training configurations, achieving best accuracies of 56.35 ± 2.67% on MEED and 23.84 ± 0.83% on DIEM-A in a subject-independent scenario. Beyond performance gains, this work also establishes new benchmark values on both datasets, providing reference results that future research on bodily emotion recognition can build on.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.