2026· Computer Modeling in Engineering & Sciences· 0 citations· 36 references
TL;DR
This study proposes a multimodal deep learning framework for Urdu emotion recognition by integrating speech and text modalities that surpasses the existing UMEDNet benchmark, demonstrating the effectiveness of transformer-based feature extraction and multimodal late fusion for Urdu emotion recognition.
Abstract
: Emotion recognition plays a crucial role in enabling intelligent human–computer interaction, yet research in low-resource languages such as Urdu remains limited, particularly in multimodal settings. This study proposes a multimodal deep learning framework for Urdu emotion recognition by integrating speech and text modalities. The approach leverages transformer-based models, namely wav2vec 2.0 for audio representation and MuRIL for text representation, combined using a late fusion strategy for classification. Experiments were conducted on the UMED dataset, consisting of 8269 multimodal instances across five emotion classes. The proposed multimodal model achieved an accuracy of 0.701 and an F1-score of 0.6915, outperforming unimodal baselines, where the audio-only and text-only models achieved accuracies of 0.6681 and 0.5085, respectively. Furthermore, the proposed approach surpasses the existing UMEDNet benchmark, demonstrating the effectiveness of transformer-based feature extraction and multimodal late fusion for Urdu emotion recognition. The results highlight the complementary nature of speech and text modalities and demonstrate that independently learned modality-specific classifiers combined through decision-level fusion can improve emotion recognition performance in low-resource languages. However, the performance improvement over alternative fusion strategies was relatively modest, indicating that more advanced multimodal interaction mechanisms may further enhance recognition performance.
Experiments show that the proposed framework outperforms unimodal baselines and existing fusion methods, indicating that the approach learns context-aware emotion representations well suited for accuracy-oriented audio–visual emotion recognition applications.
Arman Sajjadi, M. Nekou, Sayna Sarvar et al.· Signal, Image and Video Proc...· 0 citations
The Multimodal Emotion Recognition (MER) is a critical aspect in the development of human-computer interaction since it is a synthesis of non-redundant information based on the use of text, audio, and visual modalities. However, the comparative effectiveness of disparate fusion strategies and mechanisms of attention in MER has not been thoroughly studied on a single experimental paradigm. In line with this, the present study engages in a comparative analysis of four multimodal configurations, namely: early Fusion without attention, early Fusion with attention, late Fusion without attention, and late Fusion complemented by attention. Each of these configurations is evaluated on the Multimodal Emotion Lines Dataset (MELD) using harmonized training protocols to ensure a fair comparison. The methodologies use pre-trained architectures of BERT to support textual representation, Wav2Vec to support acoustic encoding, and TimesFormer to support visual streams to extract features that are modality-specific. These characteristics are then fused through the above fusion tactics. Attention modules that constitute cross attention, hierarchical attention, and self-attention modules are incorporated to enable cross-modal as well as intra-modal features interaction. The results of empirical studies have shown that there are recognizable differences between the models, where attention-enhanced schemes have better measures of performance, especially improved in terms of accuracy and weighted F1-score. However, they still have residual problems, including the imbalance of the classes that influence minority emotion categories. Such findings explain the impact created by specific fusion paradigms and attention structure on the multimodal emotion recognition performance. In turn, this research provides a methodological framework that can be used to develop more effective and understandable MER systems using a systematic and structured comparative framework.
Chintan Chatterjee, Brijesh Bhatt· International Journal of Ele...· 0 citations
: Extracting features from different data sources, such as text, images, and speech presents several challenges and integrating information from these diverse sources is complex. This study introduces an innovative hybrid method for multimodal emotion recognition using deep learning techniques, including long short-term memory (LSTM) and transformer models. These models are chosen for their ability to effectively handle sequential data in text and speech, while convolutional neural networks (CNN) are employed to analyze image data. The proposed system incorporates three distinct fusion strategies: early fusion, late fusion and hybrid fusion. In comprehensive testing, the model demonstrates high classification accuracy, achieving scores of 94.5%, 92.3%, and 93.7% on text, images, and speech respectively, using the CMU-MOSEI dataset. The model's performance, as reflected in precision, recall, and F1 scores, significantly surpasses existing state-of-the-art methods. Moreover, the model exhibits strong cross-modal synergy, contributing to overall performance in multimodal emotion recognition. This study presents a meaningful advancement in developing more accurate, scalable, and practical emotion recognition systems, with potential applications in human-computer interaction, mental health monitoring, and automated customer service.
Shailesh Kulkarni, Suhas S. Khot, Yogesh Angal· Proceedings of the 1st Inter...· 0 citations
Speech emotion recognition (SER) is a fundamental task in affective computing; however, traditional unimodal approaches often struggle to capture the complex emotional cues present in spontaneous conversational speech. Bimodal frameworks that integrate acoustic and textual information have therefore emerged to provide complementary semantic and acoustic representations. This study proposes a bimodal SER framework based on a hybrid convolutional neural network–long short-term memory (CNN–LSTM) architecture. Using the Multimodal EmotionLines Dataset (MELD), the framework combines temporal acoustic features, statistical acoustic features, and predicted textual sentiment. Experimental results indicate that the proposed model achieves reliable recognition of majority emotion classes but exhibits limited performance on underrepresented minority classes due to severe class imbalance. To better understand the contribution of each modality, feature sufficiency and feature necessity analyses were conducted. Furthermore, an evaluation of alternative fusion strategies showed that the expressive attention networks did not provide meaningful performance improvements over simple feature concatenation. These findings suggest that class imbalance, rather than fusion complexity, remains the primary limitation in conversational SER, highlighting the importance of addressing data imbalance before pursuing more sophisticated multimodal architectures.
Speech Emotion Recognition (SER) has attracted considerable attention in recent years because of its potential applications in healthcare, intelligent assistants, customer service, and human-computer interaction. However, the performance of SER systems is often affected by the imbalance of emotion classes in benchmark datasets, where minority emotions are underrepresented and difficult to classify accurately. This study reproduces the work of Sahu et al. on multimodal Speech Emotion Recognition using the Interactive Emotional Dyadic Motion Capture (IEMOCAP) dataset. Further, it proposes a Hybrid SMOTE-Ensemble Framework to improve classification performance on imbalanced data. The reproduction stage follows the original methodology by employing handcrafted acoustic features together with TF-IDF textual features under Audio-only, Text-only, and Audio+Text modalities. The enhancement stage introduces two XGBoost classifiers trained independently on the original imbalanced dataset and on a SMOTE-balanced dataset. Their prediction probabilities are combined using three ensemble strategies, namely Simple Averaging, Confidence-Based Selection, and Adaptive Weighting. In addition, the effectiveness of the proposed strategy is further investigated using a LinearSVC classifier. Experimental results demonstrate that the reproduced models achieve results that are highly consistent with those reported in the original study. Furthermore, the proposed Hybrid SMOTE-Ensemble Framework improves the recognition of minority emotion classes while maintaining balanced overall performance. The XGBoost model trained on SMOTE-balanced data achieves higher F1-score and recall than the baseline model, whereas the ensemble approaches provide additional improvements in classification performance. The LinearSVC experiments further confirm that combining oversampling and ensemble learning offers a robust solution to the class imbalance problem in Speech Emotion Recognition.
Eyad Megdadi, Mariam Aljouhi, Zainab Alblooshi et al.· International Journal of Adv...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.