Jul 2026· 2026 11th International Conference on Applying New Technology in Green Buildings (ATiGB)· pp. 1123-1128· 0 citations· 18 references
Abstract
The rapid growth of user-generated content on Vietnamese e-commerce platforms (Tiki, Google Play) has created an urgent need for accurate sentiment analysis of long documents (>256 tokens). However, existing Vietnamese Transformer models like PhoBERT are limited by the 256-token input limit, leading to a loss of context when emotional signals appear later in the text. This study presents sentiment analysis of long documents in Vietnamese. We introduce a balanced dataset of 30,000 reviews (10,000 of each sentiment type: Positive 33.33%, Neutral 33.33%, Negative 33.33%) stratified by four length groups (<256, 256-512, 512-1024, >1024 tokens). We compared PhoBERT's basic models (truncation, sliding window, hierarchy) with a Longformer model initialized from PhoBERT, expanding the context to 2048 tokens through sparse self-attention mechanisms. Experimental results showed that Longformer achieved a Macro-F1 of 0.9206 (compared to 0.8923 with truncation, +2.83%), with performance increasing positively with document length (F1=0.9975 for >1024 tokens) while the truncation method reduced performance when exceeding 512 tokens. These results confirm that explicit long-context modeling via sparse attention is essential for robust document-level sentiment analysis in Vietnamese and provide the first reproducible framework for adapting monolingual Transformer models to long-document tasks in resource-limited languages.
In the context of the digital economy, e-commerce and social media have generated massive amounts of short Chinese web texts, making the accurate extraction of sentiment information a critical requirement for market analysis and public opinion monitoring. Short texts are characterized by fragmented information expressi...
Yingying Cai, Jinliang Ma· International Conference on...· 0 citations
This study successfully proposes a Long Short-Term Memory (LSTM)-based model for automatic classification of Indonesian regional song lyrics by language, demonstrating that LSTM effectively captures sequential linguistic patterns and contextual relationships within regional languages.
Muhammad Rizky, Anandita Priatama, Aviv Yuniar Rahman et al.· Buana Information Technology...· 0 citations
A Hybrid VADER–IndoBERT framework designed to improve sentiment classification robustness on complex Indonesian texts is introduced, demonstrating the superiority of Transformer-based architectures in capturing long-range dependencies and handling ambiguous sentiment cues.
Margareta Valencia Suci Handayani, R. S. Basuki, Muljono et al.· Jurnal RESTI (Rekayasa Siste...· 0 citations
This paper presents a context-aware hybrid deep learning approach by integrating the Robustly Optimized BERT Pretraining Approach (RoBERTa) with Bidirectional Long Short-Term Memory (BiLSTM) networks to generate rich contextual word embeddings.
V. Gayatri, Rajani Rajalingam· International Journal for Re...· 0 citations
This study analyzes public sentiment in YouTube comments on felt-earthquake news in Indonesia and compares a bidirectional Long Short-Term Memory implementation (LSTM) with IndoBERT. An experimental quantitative design was used. Comments were collected through the YouTube Data API v3 using official earthquake-event ref...
Oktifar Tri Bandono, Agung Budi Susanto, Makhsun Makhsun· International Journal Of Hum...· 0 citations
The findings confirm that the combination of transformer-based models is effective for in-depth analysis of discourse in Indonesian-language political news headlines from a major Indonesian online news portal (detik.com).
Bagas Yana Prayoga, Qurrotul Aini, Fitroh Fitroh· International Journal of Int...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.