The proposed MRCF contains a Reliability-Aware Branch that estimates sample-specific modality reliability from intramodal quality cues and cross-modal semantic consistency, a Reliability-Guided Interaction Branch that uses the estimated scores to modulate cross-modal information flow, and a Reliability-Calibrated Fusion Module that integrates reliability and semantic cues for final prediction.
Abstract
Multimodal Sentiment Analysis (MSA) integrates text, audio, and vision to infer human affect, yet real-world multimodal observations are often incomplete. Existing methods for incomplete-observation MSA mainly follow two paradigms. Reconstruction-based methods recover missing information from observed modalities, while joint-representation methods learn directly from incomplete inputs. Although effective, these methods usually treat modality reliability only implicitly within representation learning or fusion design rather than modeling it explicitly. We argue that modality reliability is a central variable in incomplete-observation settings. Failure to model it explicitly gives rise to two related issues. The first is reliability mismatch, in which the affective evidence retained by each modality varies across samples and missing rates. The second is reliability propagation bias, in which messages from degraded modalities may adversely affect cross-modal interaction and predictive performance. To address these issues, we propose MRCF, a Modality Reliability-Calibrated Framework for MSA with incomplete observations. MRCF contains a Reliability-Aware Branch that estimates sample-specific modality reliability from intramodal quality cues and cross-modal semantic consistency, a Reliability-Guided Interaction Branch that uses the estimated scores to modulate cross-modal information flow, and a Reliability-Calibrated Fusion Module that integrates reliability and semantic cues for final prediction. Experiments on CMU-MOSI, CMU-MOSEI, and CH-SIMS show that MRCF achieves strong performance under standard incomplete-observation protocols. Further analyses provide evidence that explicit reliability modeling helps mitigate reliability mismatch and reliability propagation bias during interaction and fusion.
Multi-modal sentiment analysis integrates linguistic, acoustic, and visual evidence, yet the reliability of these streams varies across samples because of missing observations, masking, measurement noise, and feature corruption. This paper presents a trainable reliability-aware evidential fusion framework that estimates not only sentiment predictions but also modality-specific evidence, predictive uncertainty, observable input quality, cross-modal disagreement, and normalized sample-dependent fusion weights. Each available modality is independently encoded and processed by an evidential classification head and a quality estimation head. Availability masks enforce exact exclusion of missing streams, while estimated quality, Dirichlet uncertainty, and Jensen–Shannon disagreement jointly regulate the contribution of each observed stream. The model is optimized end-to-end using fused classification, evidential regularization, clean–corrupted consistency, reliability-calibrated cross-modal alignment, and quality regression objectives. Experiments are conducted on both CMU-MOSI and CMU-MOSEI using their official speaker-independent splits. Binary classification follows the standard non-zero protocol, in which samples with sentiment score zero are excluded from Acc-2 and binary F1 evaluation; all labeled samples are retained for seven-class accuracy, mean absolute error, and correlation. The evaluation covers complete-input, every single- and double-modality missing pattern, graded and unseen corruption, combined missing-plus-corrupted conditions, calibration, selective prediction, statistical testing, and computational efficiency. All comparative values in the main tables are identified as local controlled adaptations under the common pipeline, while selected published reference values are reported separately to prevent provenance mixing. Across both datasets, the empirical results show that the proposed method preserves competitive complete-input performance while providing larger and more consistent gains as modality availability or integrity deteriorates.
Multimodal sentiment analysis aims to infer affective states by integrating language, visual, and acoustic cues. However, real-world multimodal inputs are often incomplete or corrupted, which can weaken cross-modal complementarity and introduce misleading information into downstream fusion. Existing proxy-based methods for incomplete MSA commonly rely on one-shot proxy construction to compensate for degraded language information, but the generated proxy may be coarse or unreliable at initialization. Prematurely injecting such a proxy into multimodal reasoning can propagate initial errors and compromise sentiment prediction. To address this limitation, we propose an iterative proxy correction framework for robust incomplete MSA. Our method constructs a language-oriented proxy from non-language modalities and progressively refines it under multimodal context through gated residual correction. The corrected proxy is then adaptively fused with the observed language representation according to an estimated language reliability score, allowing the model to balance proxy-based compensation and trustworthy linguistic evidence. In addition, we introduce a stage-wise latent correction objective that uses the complete language representation as a training-time semantic anchor to stabilize the proxy refinement trajectory. Extensive experiments on MOSI, MOSEI, and SIMS under diverse missing-modality settings demonstrate that the proposed framework consistently outperforms competitive baselines and achieves robust sentiment prediction under incomplete inputs.
Zhi-Fa Geng, Subin Huang, Hao Guo et al.· 0 citations
Multimodal sentiment analysis (MSA) aims to understand human emotions by integrating heterogeneous modalities such as language and acoustic signals. Despite recent progress, two key challenges remain: (1) inter-modal inconsistency, where different modalities may convey conflicting sentiment cues, and (2) intra-modal feature ambiguity caused by noise and subtle emotional variations. To address these issues, we propose RegCal-Net, a register-augmented and self-calibrated framework for bimodal sentiment analysis. First, we introduce a Register-Augmented Self-Attention (RASA) mechanism that appends learnable register tokens along the sequence dimension to provide auxiliary global anchors for each modality. Second, we design a Self-Calibrated Fusion (SCF) module that dynamically evaluates fused feature reliability through an auxiliary score-guided gating strategy, enabling adaptive suppression of unreliable signals during multimodal integration. Extensive experiments on two widely used benchmarks, CMU-MOSI and CMU-MOSEI, together with CMU-MOSI encoder-controlled baselines and a supplementary video-subset validation, demonstrate that RegCal-Net achieves competitive performance across multiple evaluation metrics, particularly improving fine-grained sentiment classification and reducing prediction error. These results indicate that combining register-based representation stabilization with quality-aware fusion provides an effective solution for robust multimodal sentiment analysis.
Bin Xu, Wei-Yang Wang, Ao Ding· Multimedia Systems· 0 citations
MGSI first encodes audio and visual streams at short-, medium-, and long-range temporal scales, preserving both local variations and global affective trends, and applies polarity- and intensity-aware enhancement to better handle ambiguous and near-neutral samples.
Shanshan Lin, Yuesheng Wu, Chao Chen et al.· 0 citations
CHSIM, a unified MSA framework for robust text–image sentiment modeling that combines three complementary mechanisms: prediction-consistency regularization under stochastic latent masking, hierarchical cross-modal alignment, and global bidirectional sequence modeling, is proposed.
Siyang Zhang, Guangli Zhu, Qianjin Zhao et al.· Journal of Supercomputing· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.