Skip to content
Open access

Lightweight and robust audio-visual emotion recognition via multi-scale mamba temporal modeling and quality-aware expert fusion

Aug 2026 · Journal of King Saud University: Computer and Information Sciences · Vol 38 · 0 citations · 48 references

Abstract

Multimodal emotion recognition for real-time human–computer interaction requires both high recognition accuracy and low computational cost. However, existing audio-visual methods often rely on heavy Transformer-based fusion or high-capacity visual backbones, making them difficult to deploy on edge devices. Moreover, their robustness to degraded audio-visual inputs and their class-wise behavior for ambiguous emotions remain insufficiently analyzed. To address these issues, we propose a lightweight audio-visual emotion recognition framework. The visual stream uses ShuffleNet for efficient facial feature extraction, while the audio stream uses multiscale MFCCs and efficient sequence modeling to capture emotional prosody. A cross-modal auxiliary fusion module is further introduced to align audio and visual representations, and lightweight channel attention is used to emphasize emotion-relevant features. Experiments on CREMA-D and IEMO-CAP demonstrate that the proposed method achieves competitive recognition accuracy with significantly fewer parameters and lower computational cost. Additional analyses on edge-device inference, audio-visual degradation, class-wise confusion, and batch-size sensitivity further validate the efficiency and robustness of the proposed framework.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.