Aug 2026
Bidirectional joint cross-attention framework for transformer based audio–visual emotion recognition
Experiments show that the proposed framework outperforms unimodal baselines and existing fusion methods, indicating that the approach learns context-aware emotion representations well suited for accuracy-oriented audio–visual emotion recognition applications.
Arman Sajjadi, M. Nekou, Sayna Sarvar et al.
· Signal, Image and Video Proc... · 0 citations