1 paper indexed here

Fetches their full publication history.

Not the right person? Other researchers publish under this name.

Conference Jun 2026

Text-Guided Joint Interaction Network for Multimodal Sentiment Analysis

Multimodal Sentiment Analysis (MSA) aims to recognize affective information by jointly exploiting signals from multiple modalities. Among textual, acoustic, and visual inputs, the textual modality usually conveys the primary semantic information associated with sentiment, whereas the other two modalities provide complementary nonverbal evidence. Based on this observation, this paper presents the Text-Guided Joint Interaction Network (TJINet), which promotes sufficient interaction between acoustic and visual information before introducing textual guidance. First, the features of each modality are independently encoded and transformed into compact representations. Next, the Gated Cross-Attention Joint Audio-Visual Interaction (JAVI-GCA) module performs bidirectional interaction between the acoustic and visual modalities and combines their complementary information into a joint audio-visual representation. Subsequently, the Language-guided Fusion Layer employs textual features as queries to selectively retrieve sentiment-related information from the previously fused audio-visual representation. The resulting multimodal representation is finally used to generate sentiment predictions. Experiments conducted on the CMU-MOSI and CH-SIMS datasets demonstrate that TJINet delivers better overall performance than several existing advanced methods.

Jiaxing Zhou, Liujia Xu · 0 citations