A Multimodal Hybrid Framework for Enhanced Sentiment Analysis in Educational Environments: Integrating Object Detection and Contextual Captioning
Abstract
Nowadays, Developments in Sentimental Computing methods have highlighted the critical role of the student sentiment in maximizing pedagogical outcomes. However, the application of traditional and automated sentiment analysis using complex classroom situations is still complicated by the limitations of the 'conventional Convolutional Neural Networks (CNN)' and 'Recurrent Neural Networks (RNN)' to capture fine-grained emotional expressions and multi-object dynamics. This research article addresses some limitations, specifically related to 'low classification accuracy' and 'lack of contextual fusion', by a novel technique proposal of a hybrid deep learning architecture. The model presented here, combines much higher accuracy CNNs in terms of spatial features extraction and real-time object detectors. It isolates the interaction(s) between students and teachers. To overcome both the loss of situational and temporal context, we used Long Short-Term Memory (LSTM) networks combined with caption generation through Natural Language Processing (NLP). Such a (multimodal) process can allow the proposed model to decode visual environmental features into semantic features to provide meaningful improvements to the sentiment classification in the dynamic and higher occupancy contexts. The preliminary assessment has suggested that the proposed architectural fusion provides a more robust understanding of classroom affect compared to (unimodal) baselines.