Skip to content

Factorize, Reconstruct, Enhance: A Unified Framework for Multimodal Sentiment Analysis

· 0 citations · 39 references

TL;DR

FUSE-Net is proposed, a Factorized and Unified Semantic Enhancement framework that decomposes each modality into shared, specific, and noise subspaces and applies contrastive learning, an information-gain constraint, and duality constraints for structured regularization to preserve task-relevant semantics during factorization.

View source

Similar papers

Aug 2026

Aspect-guided dual-branch fusion network for multimodal aspect-based sentiment analysis

An Aspect-guided dual-branch fusion network (ADFN) to enhance sentiment prediction by incorporating external knowledge and integrating coarse and fine information is proposed, which incorporates syntactic dependency information to complement and enrich the textual semantic representations.

Bin Song, Wenjing Liu, Zhi Liang et al. · 0 citations
Open access Sep 2026

DiMoE: Disentangled Representation Learning with Mixture of Experts Fusion for Sentiment Intensity Prediction and Emotion Classification

Multimodal affective analysis benefits from combining textual, acoustic, and visual cues, yet many methods implicitly mix modality-invariant information with modality-specific factors, which can reduce robustness when modalities vary in reliability. We propose DiMoE, a representation-first framework that integrates feature disentanglement with sparse Mixture of Experts (MoE) fusion. DiMoE first decomposes each modality into shared (modality-invariant) and private (modality-specific) representations. It then fuses these factors using top-k sparse MoE routing, enabling input-adaptive expert selection. We study two routing strategies with a shared expert pool, a joint router that operates on combined shared and private factors, and separate routers that process shared and private factors independently. We evaluate DiMoE on five widely used benchmarks spanning sentiment intensity prediction and emotion classification, covering both trimodal and bimodal settings. Across benchmarks, DiMoE consistently achieves strong performance against recent competitive baselines, indicating that explicit modality disentanglement combined with Mixture of Experts fusion yields effective multimodal affective representations.

M. S. Sran, S. Ramanna, K. Kotecha · 0 citations
Open access Sep 2026

Register-augmented attention and self-calibrated fusion for robust multimodal sentiment analysis

Multimodal sentiment analysis (MSA) aims to understand human emotions by integrating heterogeneous modalities such as language and acoustic signals. Despite recent progress, two key challenges remain: (1) inter-modal inconsistency, where different modalities may convey conflicting sentiment cues, and (2) intra-modal feature ambiguity caused by noise and subtle emotional variations. To address these issues, we propose RegCal-Net, a register-augmented and self-calibrated framework for bimodal sentiment analysis. First, we introduce a Register-Augmented Self-Attention (RASA) mechanism that appends learnable register tokens along the sequence dimension to provide auxiliary global anchors for each modality. Second, we design a Self-Calibrated Fusion (SCF) module that dynamically evaluates fused feature reliability through an auxiliary score-guided gating strategy, enabling adaptive suppression of unreliable signals during multimodal integration. Extensive experiments on two widely used benchmarks, CMU-MOSI and CMU-MOSEI, together with CMU-MOSI encoder-controlled baselines and a supplementary video-subset validation, demonstrate that RegCal-Net achieves competitive performance across multiple evaluation metrics, particularly improving fine-grained sentiment classification and reducing prediction error. These results indicate that combining register-based representation stabilization with quality-aware fusion provides an effective solution for robust multimodal sentiment analysis.

Bin Xu, Wei-Yang Wang, Ao Ding · 0 citations
Open access 2026

SAMS-M: Explicit sentiment-guided alignment and multi-dimensional mutual supervision for multimodal sentiment analysis

Multimodal Sentiment Analysis leverages the fusion of heterogeneous data to achieve fine-grained emotional understanding, which finds extensive application in large-scale public opinion monitoring and data mining. However, existing methods face two key challenges: (1) cross-modal alignment suffers from redundancy and semantic drift without explicit modeling of sentiment-critical cues, inducing spurious correlations; and (2) heterogeneous representation spaces lead to imbalanced modality contributions, particularly under weak image–text correlation or sentiment inconsistency. To address these challenges, we propose an explicit sentiment-guided alignment and multi-dimensional cross-modal mutual supervisionbased model for multimodal sentiment analysis. The model primarily employs a fine-grained sentiment–saliency directed alignment mechanism, which leverages bidirectional cross-attention to couple textual sentiment cues with visual saliency, enabling precise localization of sentiment-relevant regions. Furthermore, we introduce a tripartite strong contrastive learning strategy to mitigate distribution discrepancies between heterogeneous modalities within a shared latent space, thereby enhancing cross-modal coherence and complementarity. Finally, we design a noiserobust gating-based fusion module, which, together with text augmentation and deep supervision, facilitates effective joint optimization. Experimental results show that SAMS-M obtains the best results on MVSA-Single and MSD and remains competitive on the noisier MVSA-Multiple benchmark; thus, the evidence supports strong but dataset-dependent performance rather than uniform state-of-the-art superiority.

Shi-Shu Qi, Yulei Zhang, Siyang Zhang et al. · 0 citations
Aug 2026

Multimodal sentiment analysis with multi-level representation learning and global tri-modality unified fusion.

A Global Tri-Modality Transformer (GTMT) that first performs parallel fusion of the three modalities and then conducts deep integration guided by the textual modality to achieve the cross-modal semantic alignment and correlation, significantly improving the global unified fusion effectiveness of the tri-modality information.

Zelong Li, Shu-Hua Lu, Fang Cui et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.