Skip to content

GTPF: Modality-specific graph transformer with prompt-aware fusion for multimodal sentiment analysis.

Aug 2026 · Neural Networks · Vol 205 Pt B, pp. 109498 · 0 citations · 46 references
Medicine

TL;DR

A modality-specific Graph Transformer with Prompt-aware Fusion (GTPF) framework that outperforms state-of-the-art methods across all metrics and validate the effectiveness of the Graph-Transformer co-design and prompt-aware fusion strategy.

Abstract

Multimodal Sentiment Analysis (MSA) integrates information from multiple modalities to infer sentiment. It faces two challenges: accurately modeling unimodal modalities and effectively fusing them. Existing graph neural network (GNN) based methods are limited by the over-smoothing problem caused by deep architectures, and thus typically adopt shallow structures for modality modeling. While these shallow structures can capture local information within a modality well, they struggle to capture global context. Meanwhile, current text-centric fusion methods do not explicitly align textual and non-textual modalities before fusion, which degrades fusion performance. To address these issues, we propose a modality-specific Graph Transformer with Prompt-aware Fusion (GTPF) framework. GTPF employs a Modality-Specific Graph Transformer (MSGT) architecture to explore both local and global information within modalities, enabling more accurate unimodal modeling. It also uses a Prompt-Aware Multimodal Fusion (PAMF) module that adopts prompt shifting to align textual and non-textual modalities before fusion, thereby enhancing text-centric fusion. We conduct extensive experiments on three MSA benchmarks. GTPF outperforms state-of-the-art methods across all metrics. On the CMU-MOSI dataset, GTPF achieves a relative improvement of up to 13.76% in mean absolute error (MAE) over the second-best model. On the CH-SIMS dataset, it achieves relative improvements of 6.01% in Pearson correlation (Corr) and 5.03% in five-class classification accuracy over the second-best model. These results validate the effectiveness of our Graph-Transformer co-design and prompt-aware fusion strategy.

View source

Similar papers

Conference 2026

Hierarchical Global-Local Interaction and Refinement for Multimodal Sentiment Analysis

A Modality Dropout strategy is first introduced at the input stage to alleviate over-reliance on a sin-gle modality and improve robustness and the proposed Hierarchical Global-Local Interaction and Refinement framework for Multimodal Sentiment Analysis (HGLIR) is proposed.

yuanyuan zhou · 0 citations
Aug 2026

TGHIN: text-guided hyper-modality interaction network for multimodal sentiment analysis

A Text-Guided Hyper-modality Interaction Network (TGHIN) for multimodal sentiment analysis with differentiated feature encoding strategies for each modality and a Joint-Specific Fusion (JSF) module that enables the hyper-modality representation to refocus on the core information of each modality.

Kun-Xia Wang, RenLei Ding, YiHan Ge et al. · 0 citations
Aug 2026

Aspect-guided dual-branch fusion network for multimodal aspect-based sentiment analysis

An Aspect-guided dual-branch fusion network (ADFN) to enhance sentiment prediction by incorporating external knowledge and integrating coarse and fine information is proposed, which incorporates syntactic dependency information to complement and enrich the textual semantic representations.

Bin Song, Wenjing Liu, Zhi Liang et al. · 0 citations
#small language model Preprint Aug 2026

Multi-Granularity Sentiment Integration for LLM-Based Multimodal Sentiment Analysis

MGSI first encodes audio and visual streams at short-, medium-, and long-range temporal scales, preserving both local variations and global affective trends, and applies polarity- and intensity-aware enhancement to better handle ambiguous and near-neutral samples.

Shanshan Lin, Yuesheng Wu, Chao Chen et al. · 0 citations
Aug 2026

MagXCL: enhanced multimodal adaptation gate and cross-modal contrastive learning for multimodal sentiment analysis

MagXCL, a unified framework designed to improve multimodal integration through more effective interaction between verbal and non-verbal modalities, is proposed, demonstrating the effectiveness of combining AMag with CrossCL to produce more accurate and robust multimodal sentiment predictions.

Duc-Duy Duong, Cam-Van Thi Nguyen, Duc-Trong Le · 0 citations
Jul 2026

PAMFF: Prompt alignment and multi-granularity feature fusion for multimodal aspect-based sentiment classification

A novel method called Prompt Alignment and Multi-Granularity Feature Fusion (PAMFF), which employs prompt templates to construct prompts for aspect terms and then generates soft entity pseudo-labels derived from both the prompt features and the image entity features to achieve deeper cross-modal information fusion.

Yaru Shang, Xue-Song Bai, Yimeng Zhan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.