Jul 2026· 2026 6th International Conference on Inventive Computation and Information Technologies (ICICIT)· pp. 1402-1406· 0 citations· 15 references
Abstract
With the evolution of multimedia platforms and online streaming services, the need for intelligent systems have increased to learn from heterogeneous data sources to understand human emotions and sentiments. Conventional unimodal paradigm for music emotion recognition and sentiment analysis, which relies solely on the use of single mode data (usually audio modality), has showed limited success in the task. Towards overcoming these challenges, this paper introduces an innovative cross-modal transformer based multimodal deep learning framework for combined music emotion and media sentiment analysis using synchronized multimodal representations from the CMU-MOSEI dataset. It proposes a framework which deploys Convolutional Neural Network (CNN) and Bidirectional Long Short-Term Memory (BiLSTM) models to extract music-inspired acoustic emotional features, Bidirectional Encoder Representations from Transformers (BERT) for textual sentiment representation learning, and Vision Transformer (ViT)-based feature extraction for visual emotional understanding respectively. The cross-modal transformer attention fuses heterogeneous modality representations and enhances the contextual interaction learning. Experimental evaluation shows that the proposed framework significantly outperforms traditional unimodal and multimodal approaches in terms of accuracy, precision, recall and F1-score. The proposed system is a solid and scalable solution for next generation affective multimedia analytics, intelligent recommendation systems and emotion-aware digital media applications.
A framework for learning adaptive cross-modal interactions for multimodal sentiment analysis that consistently outperforms previous methods and enhances multimodal representation capability for sentiment classification is proposed.
Chuhan Cheng, Hangcheng Wu, Jun-Qiao Wang et al.· International Conference on...· 0 citations
Speech emotion recognition (SER) is a fundamental task in affective computing; however, traditional unimodal approaches often struggle to capture the complex emotional cues present in spontaneous conversational speech. Bimodal frameworks that integrate acoustic and textual information have therefore emerged to provide complementary semantic and acoustic representations. This study proposes a bimodal SER framework based on a hybrid convolutional neural network–long short-term memory (CNN–LSTM) architecture. Using the Multimodal EmotionLines Dataset (MELD), the framework combines temporal acoustic features, statistical acoustic features, and predicted textual sentiment. Experimental results indicate that the proposed model achieves reliable recognition of majority emotion classes but exhibits limited performance on underrepresented minority classes due to severe class imbalance. To better understand the contribution of each modality, feature sufficiency and feature necessity analyses were conducted. Furthermore, an evaluation of alternative fusion strategies showed that the expressive attention networks did not provide meaningful performance improvements over simple feature concatenation. These findings suggest that class imbalance, rather than fusion complexity, remains the primary limitation in conversational SER, highlighting the importance of addressing data imbalance before pursuing more sophisticated multimodal architectures.
A hybrid Deep Auto-Encoding with Convolutional Neural Networks (DAE+CNN) is presented, a multi-modal technique based on cross-attention based multi-modal fusion model for text and visuals that outperforms single-modal sentiment analysis.
M. Yuvaraja, Dr. C. Kumuthini· Journal of Intelligent Decis...· 0 citations
In the social media ecosystem, user-generated content has gradually evolved into a multimodal form coexisting with text and speech. A single information dimension can hardly fully characterize users' opinions and emotional tendencies. To address the problems of insufficient feature representation in unimodal sentiment recognition and significant information interference in traditional fusion methods, a mid-term feature fusion multimodal analysis model combining Text Convolutional Neural Network (TextCNN) and Bidirectional Long Short-Term Memory (BiLSTM) is constructed. The model relies on a dual-branch structure to mine text semantic features and speech timing features respectively, completes cross-modal information complementary fusion through feature concatenation, and connects a fully connected layer and Softmax classifier to realize three-class sentiment discrimination. Simulation experiments based on the CMU-MOSEI dataset show that the model achieves an accuracy of 0.74 and a Macro-F1 value of 0.73, which is 2.8 percentage points higher in accuracy and 3.1 percentage points higher in Macro-F1 value than the late fusion model. Ablation experiments confirm that the text branch, speech branch and feature fusion unit all provide positive support for model performance. The mid-term fusion method of multimodal features can effectively make up for the defects of unimodal information, providing a feasible technical solution for multimodal sentiment recognition and information mining in social scenarios.
Rong Zhong· International Conference on...· 0 citations
XSentiFusionNet is an end-to-end explainable multimodal framework for audio-visual sentiment analysis that incorporates cross-modal transformer-attention, reliability-aware adaptive fusion, and XAI and showed higher robustness to noise and generalization ability on different multimodal datasets.
M. Kidwai, C. Author, Dr. Faiyaz Ahmad· Journal of Intelligent Decis...· 0 citations
MGSI first encodes audio and visual streams at short-, medium-, and long-range temporal scales, preserving both local variations and global affective trends, and applies polarity- and intensity-aware enhancement to better handle ambiguous and near-neutral samples.
Shanshan Lin, Yuesheng Wu, Chao Chen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.