This study proposes a two-stage classification approach that decomposes the multi-class classification task into two sequential binary classification stages using a BERT-based Indonesian-language transformer model (IndoBERT), and suggests that simplifying the decision space through gradual task decomposition is more effective than intervention at the loss-function level.
Abstract
The rapid growth of the Indonesian e-commerce industry has generated a large volume of customer reviews for sentiment analysis, but the data distribution often suffers from extreme class imbalance. The review dataset exhibits a 97.6% dominance of the positive class, causing the single-stage transformer model to produce high accuracy that does not fully represent classification capability. The baseline model achieves a macro-averaged F1-score of 0.599, with a neutral-class recall of 26.3%. Approaches based on loss function adjustment, such as class-balanced loss, focal loss, weighted cross-entropy, and decision-threshold adjustment, are unable to fundamentally address this issue, yielding only limited performance improvements. This study proposes a two-stage classification approach that decomposes the multi-class classification task into two sequential binary classification stages using a BERT-based Indonesian-language transformer model (IndoBERT). The first stage separates the positive class from the non-positive class, while the second stage distinguishes between the neutral and negative classes in a more balanced decision space. The proposed approach achieves a macro-averaged F1-score of 0.761, representing a 16.2% improvement over the baseline and outperforming all loss-function-based methods. These findings suggest that, under conditions of extreme class imbalance, simplifying the decision space through gradual task decomposition is more effective than intervention at the loss-function level. Furthermore, error propagation analysis and qualitative evaluations demonstrate that this approach improves sensitivity to minority classes, although challenges remain in cases involving ambiguous expressions.
The integration of the IndoBERT-BiLSTM architecture with SHAP is demonstrated to deliver accurate and explainable Indonesian sentiment analysis, which effectively bridges the gap between deep learning performance and decision transparency without compromising classification accuracy.
A. Widiyatmoko, A. Nugroho, Muhammad Nurul Firdaus· Journal of Electrical Engine...· 0 citations
Sentiment analysis has become an important task in natural language processing for understanding public opinions expressed in online reviews. However, most publicly available IMDb datasets are limited to binary sentiment labels, which restricts the ability of sentiment analysis systems to capture neutral opinions. This study proposes an efficient sentiment analysis framework that transforms the binary IMDb dataset into a three-class sentiment classification problem consisting of positive, neutral, and negative sentiments. The proposed approach integrates pseudolabeling with Parameter-Efficient Fine-Tuning (PEFT) using the Low-Rank Adaptation (LoRA) technique on the Longformer architecture. Experimental results show that the model achieves an accuracy of 77.06%, a weighted F1-score of 72.17%, and a Matthews Correlation Coefficient (MCC) of 0.6232. The results demonstrate that LoRA-based fine-tuning can significantly reduce computational requirements while maintaining competitive performance in sentiment classification tasks. These findings indicate that the proposed framework provides a practical and computationally efficient solution for large-scale sentiment analysis, particularly for environments with limited computational resources.
P. Hiskiawan, Wendy Tjung, Dustin Darmawan Isya Widjaja et al.· JRST: Jurnal Riset Sains dan...· 0 citations
A Hybrid VADER–IndoBERT framework designed to improve sentiment classification robustness on complex Indonesian texts is introduced, demonstrating the superiority of Transformer-based architectures in capturing long-range dependencies and handling ambiguous sentiment cues.
Margareta Valencia Suci Handayani, R. S. Basuki, Muljono et al.· Jurnal RESTI (Rekayasa Siste...· 0 citations
The rapid growth of social media has made it a primary channel for the public to express opinions on national strategic economic policies, including the establishment of the Danantara entity. This study aims to map public sentiment on Platform X and compare the performance of classical frequency-based architectures with transformer-based models. A common research gap in previous studies is the reliance on Bag-of-Words models, which fail to capture local context and sarcasm in informal text. A total of 9,525 tweets from the period January–May 2025 were collected via crawling and labeled using a hybrid approach combining InSet Lexicon and manual validation by experts (Cohen’s Kappa = 0.81). To address significant class imbalance (66.5% negative), SMOTE was applied to classical models. Experimental results reveal a significant performance gap: the classical TF-IDF + SVM model achieved a positive-class F1-score of only 59% due to feature distortion caused by SMOTE in the TF-IDF space, while the fine-tuned IndoBERT model substantially outperformed it with a global accuracy of 95.80% and a positive-class F1-score of 81%. These findings demonstrate that the deep transformer approach is far more robust in extracting semantics from informal Indonesian social media text, with practical implications for public policy decision-making.
S. Pradana, Etika Kartikadarma· JOURNAL OF APPLIED INFORMATI...· 0 citations
Indonesia's expanding e-commerce sector generates a growing volume of customer-written product reviews that can reveal both satisfaction and dissatisfaction. Automatically determining sentiment in these reviews is nevertheless difficult because marketplace language commonly includes informal wording, inconsistent spelling, brief statements, and domain-specific terms. This research benchmarks conventional machine learning methods for classifying the sentiment of Indonesian e-commerce reviews in the PRDECT-ID dataset. The data were obtained from Tokopedia and contain sentiment and emotion annotations. Following preprocessing, the experiment used 5,305 reviews, comprising 2,752 negative and 2,553 positive instances. The processing pipeline included case folding, text cleaning, normalization, tokenization, selective removal of stopwords, and Term Frequency-Inverse Document Frequency (TF-IDF) feature construction. Multinomial Naive Bayes, Support Vector Machine, and Random Forest were then evaluated under the same experimental configuration. The TF-IDF and Support Vector Machine combination produced the strongest results, reaching 0.9595 accuracy, 0.9594 macro-F1, and 0.9595 weighted-F1. These findings establish a reproducible reference point for sentiment classification in Indonesian e-commerce reviews.
Muhammad Fairuzabadi, Indo Intan, Sitti Suhada· JTH: Journal of Technology a...· 0 citations
This study successfully proposes a Long Short-Term Memory (LSTM)-based model for automatic classification of Indonesian regional song lyrics by language, demonstrating that LSTM effectively captures sequential linguistic patterns and contextual relationships within regional languages.
Muhammad Rizky, Anandita Priatama, Aviv Yuniar Rahman et al.· Buana Information Technology...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.