Comparative Analysis of Naïve Bayes and SMOTE-Based Long Short-Term Memory (LSTM) for Electric Vehicle Sentiment Analysis on YouTube
Abstract
The transition to electric vehicles in Indonesia has generated diverse public opinions on social media. Most previous sentiment analysis studies have tended to employ a single classification method without in-depth comparison and have overlooked the issue of extreme data imbalance, which can introduce bias into classification models. Furthermore, a research gap remains regarding the effectiveness of complex deep learning models compared with simpler statistical models when applied to medium-dimensional Indonesian-language opinion texts. This study aims to address this gap by conducting a comparative evaluation of model performance. The dataset consists of 1,512 unstructured public opinion comments collected from YouTube, distributed across three sentiment classes: 868 negative, 503 positive, and 141 neutral comments. To address the class imbalance, the Synthetic Minority Over-sampling Technique (SMOTE) was applied. The study then compares the statistical Naïve Bayes method using TF-IDF weighting with the Deep Learning Long Short-Term Memory (LSTM) method using Word Embedding. Evaluation using a Confusion Matrix demonstrates that Naïve Bayes outperformed LSTM and exhibited greater stability in sentiment classification for this dataset, achieving an accuracy of 63.04%, compared with 57.43% for the LSTM model.