Back to feed
Conference

Hierarchical Global-Local Interaction and Refinement for Multimodal Sentiment Analysis

2026 · Poster Volume 0007 The 2026 Twenty-Second International Conference on Intelligent Computing July 23-26, 2026 Toronto, Canada · 0 citations

Abstract

Multimodal Sentiment Analysis (MSA) aims to integrate text, audio, and visual modalities to achieve accurate sentiment modeling. Existing methods often rely on shallow interaction structures, making it difficult to jointly capture fine-grained local dynamics and high-level semantic dependencies. In addition, dominant modalities may suppress the learning of weaker modalities, leading to insufficient cross-modal semantic alignment. To address these issues, we pro-pose a Hierarchical Global-Local Interaction and Refinement framework for Multimodal Sentiment Analysis (HGLIR). Specifically, a Modality Dropout strategy is first introduced at the input stage to alleviate over-reliance on a sin-gle modality and improve robustness. Based on this, a Hierarchical Local Inter-action (HLI) module models multimodal sequences through a multi-layer pro-gressive structure to capture local dynamic features at different semantic levels. Within the HLI module, a Cross-Modal Synergistic Learning (CMSL) mecha-nism explicitly models cross-modal semantic consistency and gradually aligns information during interaction. Furthermore, a Global Representation Refine-ment (GRR) module introduces learnable global representations and iteratively updates them in a multi-layer structure to aggregate long-range semantic de-pendencies and form stable high-level semantic representations. Experimental results on CMU-MOSI and CMU-MOSEI demonstrate the effectiveness of the proposed framework across multiple evaluation metrics.

View source