Register-augmented attention and self-calibrated fusion for robust multimodal sentiment analysis
Abstract
Multimodal sentiment analysis (MSA) aims to understand human emotions by integrating heterogeneous modalities such as language and acoustic signals. Despite recent progress, two key challenges remain: (1) inter-modal inconsistency, where different modalities may convey conflicting sentiment cues, and (2) intra-modal feature ambiguity caused by noise and subtle emotional variations. To address these issues, we propose RegCal-Net, a register-augmented and self-calibrated framework for bimodal sentiment analysis. First, we introduce a Register-Augmented Self-Attention (RASA) mechanism that appends learnable register tokens along the sequence dimension to provide auxiliary global anchors for each modality. Second, we design a Self-Calibrated Fusion (SCF) module that dynamically evaluates fused feature reliability through an auxiliary score-guided gating strategy, enabling adaptive suppression of unreliable signals during multimodal integration. Extensive experiments on two widely used benchmarks, CMU-MOSI and CMU-MOSEI, together with CMU-MOSI encoder-controlled baselines and a supplementary video-subset validation, demonstrate that RegCal-Net achieves competitive performance across multiple evaluation metrics, particularly improving fine-grained sentiment classification and reducing prediction error. These results indicate that combining register-based representation stabilization with quality-aware fusion provides an effective solution for robust multimodal sentiment analysis.