Skip to content
Conference Open access

Multimodal Emotion Recognition: Leveraging Hybrid Fusion with Deep Learning Techniques

2025 · Proceedings of the 1st International Conference on Interdisciplinary Technology & Science Convergence (FusionX Global) · 0 citations · 15 references

Abstract

: Extracting features from different data sources, such as text, images, and speech presents several challenges and integrating information from these diverse sources is complex. This study introduces an innovative hybrid method for multimodal emotion recognition using deep learning techniques, including long short-term memory (LSTM) and transformer models. These models are chosen for their ability to effectively handle sequential data in text and speech, while convolutional neural networks (CNN) are employed to analyze image data. The proposed system incorporates three distinct fusion strategies: early fusion, late fusion and hybrid fusion. In comprehensive testing, the model demonstrates high classification accuracy, achieving scores of 94.5%, 92.3%, and 93.7% on text, images, and speech respectively, using the CMU-MOSEI dataset. The model's performance, as reflected in precision, recall, and F1 scores, significantly surpasses existing state-of-the-art methods. Moreover, the model exhibits strong cross-modal synergy, contributing to overall performance in multimodal emotion recognition. This study presents a meaningful advancement in developing more accurate, scalable, and practical emotion recognition systems, with potential applications in human-computer interaction, mental health monitoring, and automated customer service.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.