Multimodal Emotion Recognition in Urdu through Late Fusion of Fine-Tuned Speech and Text Representations
This study proposes a multimodal deep learning framework for Urdu emotion recognition by integrating speech and text modalities that surpasses the existing UMEDNet benchmark, demonstrating the effectiveness of transformer-based feature extraction and multimodal late fusion for Urdu emotion recognition.