Skip to content
Open access

WEIGHTED LOSS STRATEGY FOR BERT-BASED TWITTER SENTIMENT ANALYSIS WITHOUT SYNTHETIC OVERSAMPLING

Aug 2026 · JITK (Jurnal Ilmu Pengetahuan dan Teknologi Komputer) · 0 citations · 31 references

TL;DR

Empirical evidence is provided that, within the present experimental configuration, a properly optimized weighted loss strategy offers a viable and computationally efficient alternative to synthetic oversampling for BERT-based Twitter sentiment classification.

Abstract

The widespread adoption of ChatGPT has generated extensive public discourse across social media, necessitating robust sentiment analysis to understand collective opinions. Traditional approaches frequently employ the Synthetic Minority Over-sampling Technique (SMOTE) to address class imbalance; however, its effectiveness on short-text data remains an open question. This study develops an optimized sentiment classification model and evaluates whether competitive performance can be achieved without synthetic data augmentation. The methodology encompasses comprehensive Natural Language Processing (NLP) preprocessing and stratified data partitioning to preserve distributional characteristics. A BERT-base architecture is fine-tuned using a class-weighted Cross-Entropy loss combined with weighted random sampling, deliberately avoiding SMOTE-based oversampling. The model is trained with the AdamW optimizer (learning rate: 3 × 10⁻⁵), batch size 32, and mixed-precision training for four epochs. On 198,639 preprocessed tweets, the proposed approach achieves 93.81% accuracy, with weighted precision, recall, and F1-score of 0.9365, 0.9381, and 0.9380 respectively, outperforming the baseline by 1.75 percentage points. Per-class analysis reveals strong performance for negative (F1-score: 0.96) and positive sentiment (F1-score: 0.94), with lower neutral classification (F1-score: 0.89), attributable to the inherent heterogeneity of neutral expressions. The training-validation gap remains below 5%, consistent with adequate regularization. These findings provide empirical evidence that, within the present experimental configuration, a properly optimized weighted loss strategy offers a viable and computationally efficient alternative to synthetic oversampling for BERT-based Twitter sentiment classification. Further controlled ablation studies and statistical validation are needed to establish generalizability.

Read PDF

Similar papers

Open access Sep 2026

Comparative Analysis of Naïve Bayes and SMOTE-Based Long Short-Term Memory (LSTM) for Electric Vehicle Sentiment Analysis on YouTube

The transition to electric vehicles in Indonesia has generated diverse public opinions on social media. Most previous sentiment analysis studies have tended to employ a single classification method without in-depth comparison and have overlooked the issue of extreme data imbalance, which can introduce bias into classif...

Rayhan Gimnastiar, Fania Indah Lestari, Nadhif Daniswara Prasetyo et al. · 0 citations
Open access Aug 2026

Sentiment Analysis of Imbalanced Dataset Through Data Augmentation and Generative Annotation Using DistilBERT and Low‐Rank Fine‐Tuning

Experimental results on the Twitter US Airline Sentiment dataset demonstrate that the proposed framework achieves strong classification performance while maintaining low training complexity, highlighting the effectiveness of combining Large Language Model (LLM) based data augmentation with parameter‐efficient transform...

Hossein Nekkouei Nasrabadi, M. Moattar · 1 citation
Sep 2026

Deep Learning-based Multi-Class Sentiment Classification from Social Media Comments using LSTM Architecture

This paper presents a lightweight sentiment classification model based on Long Short-Term Memory networks, developed as a foundational text-analysis component for future multimodal emotion recognition systems, and provides a reproducible and computationally efficient baseline suitable for integration into broader multi...

Munmun Kakkar, Hemant Patidar · 0 citations
#small language model Open access Sep 2026

An Empirical Benchmarking of Traditional Machine Learning and DistilBERT-Based Zero-Shot Hierarchical Sentiment Analysis on Large-Scale Twitter Data

Text sentiment analysis of the social media text faces challenges posed by unstructured data and labori- ous human labeling for intent-driven, hierarchical classification. This work compares conventional ML models (SVM, Naïve Bayes, Logistic Regression) with contextual DL models (DistilBERT) in terms of their performan...

Bhumit Peshavariya, S. Nahar · 0 citations
Review Open access Aug 2026

AI-Driven Ensemble for Enhanced Sentiment Polarity Detection in Movie Reviews

This paper presents a weighted multi-model ensemble approach for discerning sentiment polarity in text documents, specifically consumer reviews. We address the binary classification problem of identifying positive versus negative sentiment by proposing a hybrid framework that integrates generative, discriminative, and...

Apeksha Bhuekar · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.