Aug 2026· Journal of King Saud University: Computer and Information Sciences· Vol 38· 0 citations· 60 references
TL;DR
TicCondDiffusion is the first diffusion-based framework specifically designed for MABSA, offering a new direction for this task, and achieves competitive performance compared with recent baseline approaches under the end-to-end setting.
Abstract
Multimodal Aspect-Based Sentiment Analysis (MABSA) aims to simultaneously extract aspects and predict their sentiment polarities from paired textual and visual content. Existing approaches typically formulate MABSA as either a classification or index generation task, which may limit the accuracy of aspect boundary localization, an essential component of end-to-end aspect-sentiment prediction. To address this issue, we propose TicCondDiffusion, a Tweet-Image-Caption conditioned diffusion model that reformulates MABSA as a denoising process over aspect boundary coordinates for aspect boundary localization and sentiment prediction, where the caption is an image-derived textual description used as auxiliary visual-semantic guidance. TicCondDiffusion consists of two main processes: a noising process and a conditioned denoising process. In the noising process, Gaussian noise is gradually added to ground-truth boundary coordinates over multiple timesteps to generate noisy boundary coordinates. In the denoising process, the denoising module, conditioned on Tweet-Image-Caption representations and timestep embeddings, recovers noisy boundary coordinates for boundary localization while simultaneously predicting the sentiment polarities of the corresponding aspects, with subsequent denoising steps providing further refinement. To effectively exploit multimodal information, TicCondDiffusion incorporates three specialized fusion modules: a Tweet-Image Fusion Module, a Tweet-Caption Fusion Module, and a Tweet Feature Fusion Module. Experimental results on Twitter-2015 and Twitter-2017 demonstrate that TicCondDiffusion achieves competitive performance compared with recent baseline approaches under the end-to-end setting. Experiments on the Political-Twitter dataset further show its adaptability to domains with different topical distributions. Additional analyses further validate the effectiveness and flexibility of the proposed framework. To the best of our knowledge, TicCondDiffusion is the first diffusion-based framework specifically designed for MABSA, offering a new direction for this task (The code is available on https://github.com/MEIMEIMEIMEIMEMEDA/TicCondDiffusion.).
An Aspect-guided dual-branch fusion network (ADFN) to enhance sentiment prediction by incorporating external knowledge and integrating coarse and fine information is proposed, which incorporates syntactic dependency information to complement and enrich the textual semantic representations.
Bin Song, Wenjing Liu, Zhi Liang et al.· Signal, Image and Video Proc...· 0 citations
Multimodal aspect-based sentiment analysis (MABSA) predicts the sentiment expressed toward a target aspect by jointly using textual and visual information, supporting fine-grained opinion analysis in product reviews, brand monitoring, and customer feedback. However, existing approaches remain sensitive to irrelevant visual regions, weak text–image alignment, and limited use of external knowledge. Motivated by these challenges, this study systematically evaluated two text-only large language models and four open-weight large vision-language models for aspect-level sentiment classification. The open-weight models were adapted using 4-bit quantized low-rank adaptation, while GPT-4o was assessed under zero-shot, one-shot, and five-shot in-context learning without parameter updates. Experiments were conducted on Twitter-2015, Twitter-2017, and the seven-domain MASAD dataset and evaluated using accuracy and macro-F1. Among the evaluated multimodal models, Qwen3-VL-8B-Instruct achieves the strongest performance, reaching 83.22% accuracy and 81.72% macro-F1 on Twitter-2015, 79.50% and 78.93% on Twitter-2017, and up to 99.84% and 99.83% in the Plant domain of MASAD. From a symmetry perspective, semantically aligned text–image–aspect inputs provide consistent cross-modal evidence, whereas shuffled images introduce asymmetric, symmetry-breaking information. The resulting performance degradation under shuffled-image ablation indicates that reliable aspect-level sentiment prediction depends on preserving cross-modal semantic correspondence. These findings demonstrate the effectiveness of parameter-efficient LVLM adaptation for MABSA.
Ismail Ifakir, E. Nfaoui, Abderrahim Zannou· Symmetry· 0 citations
MGSI first encodes audio and visual streams at short-, medium-, and long-range temporal scales, preserving both local variations and global affective trends, and applies polarity- and intensity-aware enhancement to better handle ambiguous and near-neutral samples.
Shanshan Lin, Yuesheng Wu, Chao Chen et al.· 0 citations
This work suggests an improved prompt-based multi-modal sentiment analysis (IPMMSA) strategy that incorporates multi-view and diversified knowledge augmentation that yields robust and expressive multimodal embedding’s to boost aspect-based sentiment analysis performance during multimodal integration with multi-modal fashion dataset.
M. Yuvaraja, C. Kumuthini· International journal of com...· 0 citations
In multimodal sentiment analysis, textual, acoustic, and visual modalities often contain redundant and noisy information. Such information increases model complexity and weakens core sentiment representations, degrading accuracy and robustness. To address this issue, we propose CLIBN, a multimodal sentiment recognition network based on contrastive learning and information bottleneck. First, we design a sentimentintensity-aware contrastive learning strategy. It constructs positive and negative pairs according to sentiment intensity distances and assigns adaptive weights to different pairs, enabling the model to capture fine-grained sentiment differences. Second, we introduce a hierarchical information bottleneck module. It treats text as the primary modality and progressively integrates complementary cues from acoustic and visual modalities, while preserving task-relevant semantics and suppressing redundant information. Experimental results on CMU-MOSI and CMU-MOSEI show that CLIBN achieves superior performance. Specifically, Acc-2 reaches 87.8% and 86.7%, and F1-Score reaches 87.8% and 86.6% on the two datasets, respectively. These results demonstrate the effectiveness of CLIBN for multimodal sentiment representation learning.
Xu Meng, Yi Zhang, Yang Li· 2026 8th International Confe...· 0 citations
Sentiment analysis is essential for several real-world applications, such as opinion mining and predicting a person's intent and personality. Most existing work aims to address challenges of sentiment analysis using normal text and images uploaded on social media. This work aims to use scene text images for sentiment analysis to assist in understanding the intentions of captured scenes. We present TSRB (Transformer-based Semantic Refinement Block), which comprises a multimodal approach and semantic gating. The proposed method constructs hierarchically fused image and text representations and then routes them through a TSRB and a learned three-way Semantic Gating module. The image branch encodes both the full meme image and text image extracted from the input image through a convolutional network with spatial attention; the text branch encodes OCR text, raw tweet text, and image captions via three independent Distil-BERT+CNN encoders and hierarchically fuses them. The resulting visual and textual embeddings are jointly refined by three stacked Transformer encoder layers within the proposed TSRB and then selectively blended by a softmax-weighted Semantic Gate that dynamically arbitrates among the post-attention, visual, and textual streams. Experiments are conducted on two standard datasets (MVSA-Single and Memotion) and compared with state-of-the-art models to demonstrate the effectiveness of the proposed method.
Soutik Mukherjee, Shivakumara Palaiahnakote, Umapada Pal et al.· International journal of pat...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.