Jul 2026· Journal of Intelligent Decision Making and Information Science· 0 citations· 30 references
TL;DR
XSentiFusionNet is an end-to-end explainable multimodal framework for audio-visual sentiment analysis that incorporates cross-modal transformer-attention, reliability-aware adaptive fusion, and XAI and showed higher robustness to noise and generalization ability on different multimodal datasets.
Abstract
The exponential growth of audio-visual content created through social media, communication platforms as well as human-computer interaction has led to a need for effective multimodal sentiment analysis. Most of the multimodal frameworks have limitations in cross-modal interaction modeling, fusion strategy adaptability, interpretability, and robustness to noise. This paper introduces XSentiFusionNet, an end-to-end explainable multimodal framework for audio-visual sentiment analysis that incorporates cross-modal transformer-attention, reliability-aware adaptive fusion, and XAI. Our approach uses CNNs, BiLSTMs, and ViTs to capture emotional features while adaptively combining acoustic and visual modalities based on their reliability levels. Model transparency is achieved by integrating SHAP, LIME, Grad-CAM, and attention mechanisms. Comprehensive experiments were performed on CMU-MOSEI, MELD, and RAVDESS benchmark datasets using comparative analysis, ablation studies, cross-dataset generalization experiments, confusion matrix analysis, explainability, and robustness to noise experiments. Our framework achieved an accuracy of 94.82%, F1-score of 94.11% and a ROC-AUC score of 96.04% on the benchmark datasets performing better than existing approaches such as CNN-LSTM, Transformer Fusion, and Multimodal BERT. Additionally, the framework showed higher robustness to noise and generalization ability on different multimodal datasets while the XAI module demonstrated interpretability by highlighting key speech and facial features used for predictions. These results show that XSentiFusionNet is a reliable and efficient framework for audio-visual sentiment analysis and can be used in real-world multimodal processing and affective computing scenarios.
Less obvious expressions like sarcasm and implicit sentiment are handled more effectively in this work, improving interpretation in multimodal sentiment analysis of social media data.
Prashant Adakane, Amit Gaikwad· international journal of eng...· 0 citations
The proposed architecture provides an effective and interpretable framework for robust multimodal sentiment analysis and offers a promising foundation for real-world affective computing applications.
B. Ankayarkanni, D. Nandini, P. Sangeetha et al.· International journal of com...· 0 citations
HAFT routes audio–visual interaction through bottleneck tokens with adaptive depth-wise gating before integrating text, jointly reducing attention cost and counteracting modality bias; the Cascaded Audio Feature Enhancement (CAFE) framework strengthens prosodic representations via multi-scale time–frequency extraction; and grouped projections, decoupled positional attention, and layer-wise parameter sharing compress the remaining overhead.
Qing Dong, Ting Lu, Xiujin Shi et al.· International journal of sof...· 0 citations
A hybrid Deep Auto-Encoding with Convolutional Neural Networks (DAE+CNN) is presented, a multi-modal technique based on cross-attention based multi-modal fusion model for text and visuals that outperforms single-modal sentiment analysis.
M. Yuvaraja, Dr. C. Kumuthini· Journal of Intelligent Decis...· 0 citations
MagXCL, a unified framework designed to improve multimodal integration through more effective interaction between verbal and non-verbal modalities, is proposed, demonstrating the effectiveness of combining AMag with CrossCL to produce more accurate and robust multimodal sentiment predictions.
A framework for learning adaptive cross-modal interactions for multimodal sentiment analysis that consistently outperforms previous methods and enhances multimodal representation capability for sentiment classification is proposed.
Chuhan Cheng, Hangcheng Wu, Jun-Qiao Wang et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.