Skip to content
Open access

Cross-Modal Attention Mechanisms in Transformer Networks for Real-Time Emotion Detection in Video Conferencing

Jul 2026 · Journal of Intelligent Decision Making and Information Science · 0 citations · 22 references

Abstract

Emotion detection in video conferencing requires models that can interpret facial-expression dynamics and speech cues under variable communication conditions. This study develops a reliability-aware RA-Gated Cross-Modal Fusion framework for real-time audio-visual emotion recognition using the RAVDESS full audio-visual speech subset. The methodology includes synchronized audio-visual preprocessing, face-frame extraction, log-mel spectrogram generation, unimodal baseline training, cross-modal feature extraction, and reliability-aware gated fusion. Two evaluation protocols were used: a stratified-random split for main benchmark assessment and an actor-independent split for speaker-disjoint generalization. The RA-Gated Cross-Modal Fusion model achieved the strongest stratified-random performance, reaching 93.55% accuracy, 93.14% macro-F1, and 93.59% weighted-F1, outperforming audio-only, visual-only, and late-fusion baselines. Class-wise analysis further showed strong recognition of angry, fearful, disgust, happy, and improved performance for sad and surprised expressions. The actor-independent evaluation showed lower performance across all models, with tuned late fusion achieving the highest speaker-disjoint result and RA-Gated fusion improving over unimodal baselines. These findings indicate that reliability-aware fusion improves benchmark emotion recognition, while the 21.50 ms per-sample RA-Gated latency on an NVIDIA GeForce RTX 3060 supports real-time-oriented applicability. Future work should test larger naturalistic conferencing datasets and strengthen speaker-invariant streaming transformers.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.