Skip to content
Open access

IndoBERT-Based Sentiment Analysis of Indonesian Social Media Discourse on AI-Generated Images

Jul 2026 · SinkrOn · Vol 10, pp. 1464-1475 · 0 citations · 30 references

TL;DR

This study establishes a robust empirical baseline for Indonesian sentiment analysis, proving transformer architectures superior for nuanced public opinion mining by fine-tuning IndoBERT and benchmarking it against classical machine learning classifiers for classifying social media sentiment.

Abstract

The rapid emergence of generative artificial intelligence has disrupted creative ecosystems, prompting widespread discourse across Indonesian social media. However, the exact sentiment structure of this public reaction remains empirically unmapped due to the contextual complexities of informal language. The objective of this research is to evaluate the efficacy of contextual language models by fine-tuning IndoBERT and benchmarking it against classical machine learning classifiers—including Complement Naive Bayes, Logistic Regression, and Support Vector Machine—for classifying social media sentiment. A multi-platform dataset comprising 2,981 Indonesian-language posts from X, Reddit, and YouTube was collected and manually annotated into positive, neutral, and negative classes. To address inherent class imbalance, Synthetic Minority Oversampling Technique was applied to classical models, while class-weighted loss and Masked Language Modeling augmentation were utilized for IndoBERT. Performance was evaluated using macro-averaged F1-score across five repeated stratified random splits. IndoBERT achieved a mean macro-F1 of 0.7131 ± 0.0180, outperforming the best classical baseline by approximately 0.12, demonstrating a pronounced advantage in resolving ambiguous neutral discourse. Negative sentiment heavily dominated the corpus at 61.8%, reflecting a prevailing critical stance toward AI-generated imagery concerning ethical and copyright issues. Furthermore, evaluation variance across random seeds exceeded variance from augmentation strategies, indicating test set composition is a major performance determinant. In conclusion, this study establishes a robust empirical baseline for Indonesian sentiment analysis, proving transformer architectures superior for nuanced public opinion mining.

Read PDF

Similar papers

Review Open access Jul 2026

A Hybrid VADER–IndoBERT Framework for Robust Sentiment Analysis of Long and Ambiguous Indonesian Texts

A Hybrid VADER–IndoBERT framework designed to improve sentiment classification robustness on complex Indonesian texts is introduced, demonstrating the superiority of Transformer-based architectures in capturing long-range dependencies and handling ambiguous sentiment cues.

Margareta Valencia Suci Handayani, R. S. Basuki, Muljono et al. · 0 citations
Open access Aug 2026

A Comparison of Classical Machine Learning and IndoBERT on Sentiment Analysis of Danantara Program in X

The rapid growth of social media has made it a primary channel for the public to express opinions on national strategic economic policies, including the establishment of the Danantara entity. This study aims to map public sentiment on Platform X and compare the performance of classical frequency-based architectures wit...

S. Pradana, Etika Kartikadarma · 0 citations
Open access Jul 2026

Evaluating Audience Perception in Indonesian Animation: Comparative Aspect-Based Sentiment Analysis Using Random Forest and IndoBERT

The rapid expansion of the Indonesian animation industry has sparked vibrant public discourse on social media, yet existing research primarily focuses on binary or overall sentiment rather than fine-grained aspect-level evaluations. This study presents an aspect-based sentiment analysis (ABSA) comparing a traditional m...

Eko Rachmat Slamet .H Saputra, A. Frobenius · 0 citations
Open access Aug 2026

Linguistically Informed Machine Learning for Gujarati–English Code-Mixed Sentiment Classification: A Comparative Study of Feature Fusion Strategies

Overall, this work demonstrates that incorporating explicit linguistic information, including language identity, sentiment polarity, and intensifier information, improves sentiment classification of Gujarati–English code-mixed text.

Chirag D. Shah, Shailesh A. Chaudhari · 0 citations
Open access Jul 2026

PANCASILA-BASED ASPECT CATEGORY SENTIMENT ANALYSIS FOR DETECTING NEGATIVE CONTENT IN INDONESIAN CODE-MIXED SOCIAL MEDIA

Hate speech and value-violating content on Indonesian social media, compounded by code-mixed language, threaten social cohesion. This study proposes a Pancasila-based Aspect Category Sentiment Analysis framework grounded in Indonesia’s five foundational values: Divinity, Humanity, Unity, Democracy, and Social Justice....

Stefani Tasya Hallatu, R. Anggraini, Adhatus Solichah Ahmadiyah · 0 citations
Open access Jul 2026

Explainable AI For Social Media Opinion Analysis Using Efficient Language Modelling

Social-media sentiment classification is difficult because posts are short, informal, context-dependent, and unevenly distributed across classes. It was observed that five TF-IDF classifiers (Logistic Regression, Linear SVM, Multinomial Naïve Bayes, Random Forest, and Gradient Boosting) were compared with the compact D...

Hadeel Saed, Mauro Alvarez, Dr.Ferhat Atik et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.