Skip to content
Open access

MULTI-MODAL TRANSFORMER ARCHITECTURE WITH CROSS-ATTENTION FUSION FOR ROBUST AUDIO-VISUAL SENTIMENT ANALYSIS

Jul 2026 · International journal of computer information systems and industrial management applications · 0 citations

Abstract

Multimodal sentiment analysis (MSA) has gained significant attention due to its ability to integrate heterogeneous information from audio, visual, and textual modalities. However, existing transformer-based fusion methods often suffer from reduced robustness when one or more modalities are corrupted or partially unavailable. This paper presents a Multi-Modal Transformer Architecture with Cross-Attention Fusion (MMT-CAF) for robust audio-visual sentiment analysis. The proposed framework combines modality-specific transformer encoders, bidirectional cross-attention, and a reliability-aware fusion mechanism that dynamically adjusts the contribution of each modality according to its estimated reliability. The framework was evaluated on the CMU-MOSI and CMU-MOSEI benchmark datasets and compared with representative transformer-based methods, including Adaptive Modality Weighting, RAFT, and CITN-DAF. Experimental results demonstrate that MMT-CAF achieves superior sentiment classification performance while maintaining higher robustness under noisy audio, visual occlusion, and missing-modality scenarios. Ablation studies further confirm the effectiveness of the proposed cross-attention and reliability-aware fusion modules in improving multimodal representation learning. The proposed architecture provides an effective and interpretable framework for robust multimodal sentiment analysis and offers a promising foundation for real-world affective computing applications.

Read PDF