Skip to content
Open access

LSTM-Based Classification of Indonesian Regional Song Lyrics by Language

Jul 2026 · Buana Information Technology and Computer Sciences (BIT and CS) · 0 citations · 14 references

TL;DR

This study successfully proposes a Long Short-Term Memory (LSTM)-based model for automatic classification of Indonesian regional song lyrics by language, demonstrating that LSTM effectively captures sequential linguistic patterns and contextual relationships within regional languages.

Abstract

This study successfully proposes a Long Short-Term Memory (LSTM)-based model for automatic classification of Indonesian regional song lyrics by language. Unlike prior works that often focus on sentiment analysis or use unbalanced datasets, this research utilizes a balanced dataset consisting of 2,500 lyric segments from five regional languages: Javanese, Sundanese, Batak, Minangkabau, and Banjarese. A comprehensive preprocessing pipeline is applied, including case folding, text cleaning, tokenization, stopword removal, stemming, sequence padding, and label encoding to transform textual data into numerical representations. The model is evaluated using 5-fold cross-validation to ensure robustness and generalization across different data partitions. Experimental results show that the proposed model achieves an accuracy of 95.24%, precision of 95.36%, recall of 95.24%, and F1-score of 95.26%, indicating strong and consistent performance. These findings demonstrate that LSTM effectively captures sequential linguistic patterns and contextual relationships within regional languages, enabling accurate classification despite similarities in vocabulary and structure. Furthermore, this study contributes to the advancement of natural language processing for low-resource languages and highlights the potential of deep learning approaches in supporting the digital preservation and automatic organization of Indonesian regional cultural content.

Read PDF

Similar papers

Open access Jul 2026

Implementation of a Bi-LSTM Model for Automatic Text Classification of Mathematics, Science, and Indonesian Language Questions

This study aims to implement the Bidirectional Long Short-Term Memory (Bi-LSTM) model for automatic text classification of Mathematics, Natural Sciences (IPA), and Indonesian language questions to support efficient question grouping in digital education systems. The dataset used consists of 2,718 questions, which are evenly distributed across three subject categories. The research stages include text preprocessing, tokenization and padding, splitting the dataset into training and testing sets, designing the Bi-LSTM model architecture, and conducting training and evaluation using accuracy, precision, recall, and F1-score metrics. The results show that the Bi-LSTM model achieves an accuracy of 97% on the test data, with an average F1-score of 0.97. The confusion matrix analysis indicates that most predictions are correctly classified with a relatively low misclassification rate across categories. Based on these results, it can be concluded that the Bi-LSTM model is effective for automatic text classification of educational questions and has strong potential for further development in technology-based question grouping systems.

Mochamad soffan Muslim, Aviv Yuniar Rahman, Rangga Pahlevi · 0 citations
Open access Aug 2026

Implementation of the IndoBERT-LSTM Model for Indonesian Sentiment Analysis withan Explainable AI Approach Using SHAP

This study aims to develop an Indonesian sentiment analysis model that achieves high classification performance while providing post-hoc explanations of prediction results.The study utilized a quantitative experimental approach using the IndoNLU SmSA dataset, comprising 11,000 training and 1,260 validation samples across positive, neutral, and negative categories. The proposed model integrates an IndoBERT contextual representation generator with a Bidirectional Long Short-Term Memory (BiLSTM) network to model sequential relationships. Furthermore, SHapley Additive exPlanations (SHAP) are applied to provide post-hoc interpretations of the model's predictions by identifying individual token contributions. The IndoBERT-BiLSTM model achieved an accuracy of 92.78%, a Macro F1-score of 0.9013, and a Macro Average AUC of 0.976, outperforming standard LSTM and fine-tuned IndoBERT baselines. However, learning curve analysis indicated mild overfitting during the training process. SHAP visualizations successfully explained the token-level contributions to the classification decisions, providing transparency into the model's reasoning. This study demonstrates the integration of the IndoBERT-BiLSTM architecture with SHAP to deliver accurate and explainable Indonesian sentiment analysis. The approach effectively bridges the gap between deep learning performance and decision transparency without compromising classification accuracy.

A. Widiyatmoko, A. Nugroho, Muhammad Nurul Firdaus · 0 citations
Open access Aug 2026

Arabic Plagiarism Detection Using Word2Vec-Based Semantic Features and Random Forest Classification on the ExAraPlagDet Dataset

The findings underscore the potential of advanced NLP techniques to overcome language-specific challenges, providing a foundation for future research in multilingual plagiarism detection and enhancing the development of tools for other languages facing similar challenges.

Hanan Mohammed Fawzy, Ahmad Salah, Heba El-Fiqi et al. · 0 citations
Review Open access Jul 2026

A Hybrid VADER–IndoBERT Framework for Robust Sentiment Analysis of Long and Ambiguous Indonesian Texts

The rapid expansion of digital learning platforms has increased the reliance on user-generated reviews for service evaluation and quality monitoring. However, sentiment analysis of Indonesian reviews remains challenging due to the prevalence of long sentences, mixed sentiments, and ambiguous linguistic expressions. This study introduces a Hybrid VADER–IndoBERT framework designed to improve sentiment classification robustness on complex Indonesian texts. A dataset of 4,904 Ruangguru application reviews was collected through web scraping and processed using a hybrid pipeline consisting of preprocessing, translation-based silver-standard sentiment labeling with VADER, and class balancing via Random Oversampling (ROS). The IndoBERT classifier was evaluated against a Bidirectional Long Short-Term Memory (BiLSTM) baseline. Experimental results show that IndoBERT achieved 90.9% accuracy, outperforming BiLSTM at 86.4%, demonstrating the superiority of Transformer-based architectures in capturing long-range dependencies and handling ambiguous sentiment cues. These findings highlight the effectiveness of integrating lexicon-based and Transformer-based approaches to achieve more robust sentiment analysis on linguistically complex Indonesian texts.

Margareta Valencia Suci Handayani, R. S. Basuki, Muljono et al. · 0 citations
Open access 2026

Arabic News Text Classification Using Deep Learning Models with Dynamic N-grams

The complexity and morphological richness of the Arabic language pose significant challenges in natural language processing (NLP), including issues with contextual understanding and feature extraction. Traditional deep learning architectures such as CNNs, LSTMs, and GRUs often struggle to effectively model these linguistic intricacies, limiting their performance on Arabic text analysis tasks. To address these limitations, the study integrates parallel multi-kernel word-level convolutional features into conventional and hybrid deep learning models. The convolutional windows capture short contextual relationships among neighboring word tokens, while recurrent components model longer sequential evidence; the framework does not directly analyze roots, affixes, or other within-word morphological structures. These enhancements are integrated into CNN, LSTM, GRU, and hybrid architectures such as LSTM-CNN and GRU-CNN. A comprehensive evaluation was conducted across varying learning rates to assess the impact of the enhanced configurations on model performance. The results indicate competitive performance within the evaluated architectures and dataset variants, although the magnitude of improvement depends on the model and learning rate. Under their best settings, the Dynamic N-gram LSTM-CNN achieved an accuracy of 93.32%, while the Dynamic N-gram LSTM achieved 93.57%. Because previously published studies use different corpora, class configurations, preprocessing pipelines, and evaluation protocols, these results are not presented as evidence of state-of-the-art superiority. Instead, the study provides a systematic within-study assessment of model sensitivity to architecture, preprocessing, and learning-rate selection. Future directions include character- and subword-level modeling, transformer-based architectures, and domain-specific tasks such as sentiment analysis and information retrieval.

Ahmed I.Taloba, George Samy Rady, Khaled F. Hussain · 0 citations
Open access 2026

A Large-Scale Vietnamese News Dataset for Text Classification: Construction and Evaluation

The Vietnamese language presents distinctive natural language processing (NLP) challenges, which are compounded by a critical shortage of standard benchmark datasets. To address this gap, this paper introduces BN-VN3S, a large-scale Vietnamese news dataset containing 946,696 final processed articles derived from 1,042,295 articles initially collected from three major Vietnamese news publishers: VnExpress, VietNamNet, and Dân TrÍ. Spanning an extended publication period from 2019 to 2025 across 10 topical categories, this dataset provides unprecedented diversity. Utilizing this resource, we conduct a comprehensive empirical evaluation of ten different models across three paradigms: Traditional Machine Learning (Naïve Bayes, Logistic Regression, LinearSVC, SGDClassifier), Deep Learning (TextCNN, BiGRU, TextRCNN), and Transformers (mBERT, XLM-R, PhoBERT). Our study analyzes classification performance through four critical dimensions: model architecture, temporal data shift, source-origin bias, and training data scale. Experimental results surprisingly reveal that deep learning and traditional models surpass Transformers in overall performance; BiGRU achieved the highest macro F1-score of 92.42%, closely followed by LinearSVC at 92.33%, whereas PhoBERT reached 87.98% and mBERT lagged at 78.73%. However, under temporal distribution shifts evaluated on 2023–2025 data, Transformers particularly PhoBERT demonstrate superior robustness and maintain the most stable performance. Furthermore, we find that models are highly sensitive to source bias; BiGRU suffered a substantial performance drop of up to 9.71 F1 points during cross-source evaluation, while Naïve Bayes and mBERT were significantly more resilient. Finally, the data-scale analysis reveals distinct learning behaviors across model families: transformer models benefit increasingly from larger training sets, whereas the strongest traditional and deep learning baselines remain competitive throughout the evaluated range. No consistent crossover is observed under the present experimental configuration, suggesting differences in sample efficiency rather than the general superiority of any model family. Taken together, these findings provide actionable insights and practical recommendations for optimizing Vietnamese news classification systems in real-world environments.

B. Chau, D. Duong, Phuoc Tran · 0 citations