Multimodal Emotion Recognition: Leveraging Hybrid Fusion with Deep Learning Techniques
Abstract
: Extracting features from different data sources, such as text, images, and speech presents several challenges and integrating information from these diverse sources is complex. This study introduces an innovative hybrid method for multimodal emotion recognition using deep learning techniques, including long short-term memory (LSTM) and transformer models. These models are chosen for their ability to effectively handle sequential data in text and speech, while convolutional neural networks (CNN) are employed to analyze image data. The proposed system incorporates three distinct fusion strategies: early fusion, late fusion and hybrid fusion. In comprehensive testing, the model demonstrates high classification accuracy, achieving scores of 94.5%, 92.3%, and 93.7% on text, images, and speech respectively, using the CMU-MOSEI dataset. The model's performance, as reflected in precision, recall, and F1 scores, significantly surpasses existing state-of-the-art methods. Moreover, the model exhibits strong cross-modal synergy, contributing to overall performance in multimodal emotion recognition. This study presents a meaningful advancement in developing more accurate, scalable, and practical emotion recognition systems, with potential applications in human-computer interaction, mental health monitoring, and automated customer service.