Skip to content

Author

Shivakumara Palaiahnakote

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

MAFN: Multimodal Attention Fusion Network for Emotion Recognition

Emotion recognition from physiological signals is a rapidly advancing area within affective computing and human-computer interfacing. This study presents a novel technique for emotion recognition that leverages Electroencephalogram (EEG) and Electrocardiogram (ECG) signals. It is observed that as emotions change, the patterns of EEG and ECG signals also change. This observation inspired us to propose a new Multimodal Attention Fusion Network (MAFN). This MAFN integrates Bidirectional Long Short-Term Memory (BiLSTM) and self-attention mechanisms to extract effective features for emotion classification. In this work, the adapted BiLSTM extracts spatial and temporal features, while a self-attention network extracts contextual features to improve the classification performance. To evaluate the model's performance, three benchmark datasets, DREAMER, AMIGOS, and Multimodal, are used to validate the proposed and existing models with a 5-fold nested cross-validation approach. Extensive experiments and analyses across all three datasets confirm the effectiveness of this approach in emotion classification. A comparative study of the proposed model with the state-of-the-art emotion recognition models shows that our work consistently surpasses state-of-the-art models on different benchmark datasets in terms of classification rate.

S. Gornale, Shivakumara Palaiahnakote, Amruta Unki et al. · 0 citations
Aug 2026

TSRB: Transformer-based Semantic Refinement Block for Sentiment Analysis using Scene Text Images

Sentiment analysis is essential for several real-world applications, such as opinion mining and predicting a person's intent and personality. Most existing work aims to address challenges of sentiment analysis using normal text and images uploaded on social media. This work aims to use scene text images for sentiment analysis to assist in understanding the intentions of captured scenes. We present TSRB (Transformer-based Semantic Refinement Block), which comprises a multimodal approach and semantic gating. The proposed method constructs hierarchically fused image and text representations and then routes them through a TSRB and a learned three-way Semantic Gating module. The image branch encodes both the full meme image and text image extracted from the input image through a convolutional network with spatial attention; the text branch encodes OCR text, raw tweet text, and image captions via three independent Distil-BERT+CNN encoders and hierarchically fuses them. The resulting visual and textual embeddings are jointly refined by three stacked Transformer encoder layers within the proposed TSRB and then selectively blended by a softmax-weighted Semantic Gate that dynamically arbitrates among the post-attention, visual, and textual streams. Experiments are conducted on two standard datasets (MVSA-Single and Memotion) and compared with state-of-the-art models to demonstrate the effectiveness of the proposed method.

Soutik Mukherjee, Shivakumara Palaiahnakote, Umapada Pal et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.