Jul 2026· 2026 8th International Conference on Electronics and Communication, Network and Computer Technology (ECNCT)· pp. 841-844· 0 citations· 12 references
Abstract
In multimodal sentiment analysis, textual, acoustic, and visual modalities often contain redundant and noisy information. Such information increases model complexity and weakens core sentiment representations, degrading accuracy and robustness. To address this issue, we propose CLIBN, a multimodal sentiment recognition network based on contrastive learning and information bottleneck. First, we design a sentimentintensity-aware contrastive learning strategy. It constructs positive and negative pairs according to sentiment intensity distances and assigns adaptive weights to different pairs, enabling the model to capture fine-grained sentiment differences. Second, we introduce a hierarchical information bottleneck module. It treats text as the primary modality and progressively integrates complementary cues from acoustic and visual modalities, while preserving task-relevant semantics and suppressing redundant information. Experimental results on CMU-MOSI and CMU-MOSEI show that CLIBN achieves superior performance. Specifically, Acc-2 reaches 87.8% and 86.7%, and F1-Score reaches 87.8% and 86.6% on the two datasets, respectively. These results demonstrate the effectiveness of CLIBN for multimodal sentiment representation learning.
MGSI first encodes audio and visual streams at short-, medium-, and long-range temporal scales, preserving both local variations and global affective trends, and applies polarity- and intensity-aware enhancement to better handle ambiguous and near-neutral samples.
Shanshan Lin, Yuesheng Wu, Chao Chen et al.· 0 citations
An Aspect-guided dual-branch fusion network (ADFN) to enhance sentiment prediction by incorporating external knowledge and integrating coarse and fine information is proposed, which incorporates syntactic dependency information to complement and enrich the textual semantic representations.
Bin Song, Wenjing Liu, Zhi Liang et al.· Signal, Image and Video Proc...· 0 citations
Consistency-Aware Gated Fusion (CAGF), a lightweight and fusion module tailored to Mamba-based architectures that achieves state-of-the-art performance, outperforming strong multimodal baselines such as CLIP, MISA, DLF, AoM, and SFTTR, while remaining more efficient and interpretable.
Jian Hu· Poster Volume 0008 The 2026...· 0 citations
A framework for learning adaptive cross-modal interactions for multimodal sentiment analysis that consistently outperforms previous methods and enhances multimodal representation capability for sentiment classification is proposed.
Chuhan Cheng, Hangcheng Wu, Jun-Qiao Wang et al.· International Conference on...· 0 citations
: Multimodal sentiment analysis (MSA) has made significant progress in integrating heterogeneous information from text, speech, and vision. However, real-world multimodal data often suffer from modality noise, semantic inconsistency, and incomplete modality information, which can weaken cross-modal fusion and reduce the reliability of sentiment prediction. To address these challenges, this paper proposes RUAL, a robust uncertainty-aware learning framework for multimodal sentiment analysis. Specifically, RUAL first employs a Gathered Multi-Head Attention Pooling (GMHA) module to aggregate intra-modal features and estimate modality uncertainty based on attention entropy. Then, an Uncertainty-Aware Cross-Modal Coupled Layer (UACCL) is introduced to dynamically regulate cross-modal residual fusion according to sample confidence, thereby reducing the negative influence of unreliable modalities on fused representations. In addition, uncertainty-weighted learning and uncertainty-guided self-distillation (UWL and U-SD) are jointly integrated through an optimization strategy to further improve training stability and generalization in complex scenarios. Experimental results on CMU-MOSI, CMU-MOSEI, and MVSA-Single demonstrate that RUAL achieves strong overall performance and maintains stable prediction results under missing-modality and Gaussian-noise conditions, validating the effectiveness and robustness of the proposed framework for multimodal sentiment analysis.
DualScope is proposed, a novel model that combines a global-local fusion strategy with bidirectional image-text generation for semantically consistent data augmentation and introduces both label contrastive learning and data contrastive learning to align heterogeneous modalities and enhance model robustness.
Bing Zhang, Junteng Wang, Bin Sun et al.· Memetic Computing· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.