Skip to content
Open access

Sentiment Analysis of Imbalanced Dataset Through Data Augmentation and Generative Annotation Using DistilBERT and Low‐Rank Fine‐Tuning

Aug 2026 · Applied AI Letters · Vol 7 · 1 citation · 35 references

TL;DR

Experimental results on the Twitter US Airline Sentiment dataset demonstrate that the proposed framework achieves strong classification performance while maintaining low training complexity, highlighting the effectiveness of combining Large Language Model (LLM) based data augmentation with parameter‐efficient transformer adaptation.

Abstract

Sentiment analysis on social media data often suffers from severe class imbalance, which can negatively affect the performance of machine learning models. In this paper, we propose a framework that leverages large language models and lightweight transformer fine tuning to improve sentiment classification on imbalanced datasets. First, GPT‐4, a multimodal large language model, is used to generate synthetic tweets through paraphrasing and back translation, with Italian serving as an intermediate language, to increase data diversity. In addition, GPT‐4 is employed to annotate tweets with positive reasons by generating semantic counterparts to the 10 predefined negative categories in the Twitter US Airline Sentiment dataset. This process enables the creation of meaningful positive annotations derived from existing category structures, thereby improving dataset balance and interpretability. The augmented data are then encoded using DistilBERT to obtain sentence embeddings, while low rank adaptation (LoRA) is applied for efficient fine tuning with reduced computational cost. Finally, a SoftMax classifier is used to predict sentiment labels (positive, neutral, and negative). Experimental results on the Twitter US Airline Sentiment dataset—evaluated rigorously using a held‐out test set and 10‐fold cross‐validation—demonstrate that the proposed framework achieves strong classification performance while maintaining low training complexity, highlighting the effectiveness of combining Large Language Model (LLM) based data augmentation with parameter‐efficient transformer adaptation.

Read PDF

Similar papers

Open access 2026

A Hybrid Framework for Large-Scale Tweet Sentiment Analysis Using Classical Machine Learning, Transformer Models, and Uncertainty Estimation

The work provides a reproducible, explainable, operationally applicable model of sentiment analysis in operationally sensitive, high-stakes Twitter sentiment analysis, and validate the hypothesis that hybrid stacking is an effective method for leveraging the complementary nature of lexical and contextual representation...

D. Abate, Nilay Mistry · 0 citations
Sep 2026

Deep Learning-based Multi-Class Sentiment Classification from Social Media Comments using LSTM Architecture

This paper presents a lightweight sentiment classification model based on Long Short-Term Memory networks, developed as a foundational text-analysis component for future multimodal emotion recognition systems, and provides a reproducible and computationally efficient baseline suitable for integration into broader multi...

Munmun Kakkar, Hemant Patidar · 0 citations
Open access Aug 2026

WEIGHTED LOSS STRATEGY FOR BERT-BASED TWITTER SENTIMENT ANALYSIS WITHOUT SYNTHETIC OVERSAMPLING

Empirical evidence is provided that, within the present experimental configuration, a properly optimized weighted loss strategy offers a viable and computationally efficient alternative to synthetic oversampling for BERT-based Twitter sentiment classification.

T. Siallagan, R. Winanjaya, Juni Ismail · 0 citations
Open access Aug 2026

Performance Evaluation of Word2Vec and FastText Embeddings in a CNN-BiLSTM Model for Sentiment Classification of the LPDP Alumni Controversy

This study aims to analyze public sentiment toward the LPDP alumni controversy on social media using a deep learning approach. The research data consist of YouTube user comments related to the LPDP issue, which were processed through text preprocessing and automatically labeled using IndoBERT into three sentiment class...

Dwi Erzalianti, Joice Junansi Tandirerung, C. Suhaeni et al. · 0 citations
#small language model Open access Sep 2026

An Empirical Benchmarking of Traditional Machine Learning and DistilBERT-Based Zero-Shot Hierarchical Sentiment Analysis on Large-Scale Twitter Data

Text sentiment analysis of the social media text faces challenges posed by unstructured data and labori- ous human labeling for intent-driven, hierarchical classification. This work compares conventional ML models (SVM, Naïve Bayes, Logistic Regression) with contextual DL models (DistilBERT) in terms of their performan...

Bhumit Peshavariya, S. Nahar · 0 citations
Sep 2026

Large Language Model-Driven Hard Negatives Generation for Contrastive Aspect-Based Sentiment Analysis

While sentiment analysis has advanced significantly, fine-grained sentiment classification such as aspect-based sentiment analysis (ABSA), continues to present challenges. These difficulties primarily stem from data scarcity and the inherent complexities of identifying sentiments specific to different aspects within...

Ling-Ling Xu, Hao-Ran Xie, S. Qin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.