Skip to content
Review Open access

Sentiment Analysis of Textual Data in Social Networks: A Survey of Existing Approaches

Aug 2026 · “Kibertəhlükəsizlik və rəqəmsal kriminalistikanın aktual problemləri” respublika elmi-praktiki konfransı · 0 citations

TL;DR

Examination of sentiment analysis methods applied to social media text data, covering lexicon-based, machine learning, and deep learning approaches, including transformerbased architectures, as well as widely used datasets, shows how sentiment analysis can be applied to the detection of threats that exploit human emotions.

Abstract

The growth of social media platforms results in billions of user-generated messages daily, making automated text analysis critical. The global impact of disinformation, the evolution of cyber threats toward psychological tactics, and the growing political and commercial value of public opinion have made sentiment analysis an important and relevant area of research. Although existing reviews have examined sentiment analysis approaches in terms of methods and application areas, few have evaluated these approaches in the context of cyber threat detection. This paper examines sentiment analysis methods applied to social media text data, covering lexicon-based, machine learning, and deep learning approaches, including transformerbased architectures, as well as widely used datasets. The paper also discusses how sentiment analysis can be applied to the detection of threats that exploit human emotions, including phishing, disinformation, and social engineering. To support this, an empirical analysis of large language model performance is conducted, measuring their ability to detect emotionally manipulative content. The purpose of this paper is to provide readers with an objective understanding of sentiment analysis and its role as a defense against socially engineered cyber threats.

Read PDF

Similar papers

Open access Jul 2026

Efficient Sentiment, Emotion, and Toxic Speech Analysis on Social Media Using Lightweight LLMs: A Practical Approach for Digital Transformation

Social media generates massive amounts of unstructured, multilingual textual data that contain many sentiments, emotions, and opinions from users that provide insights for enterprises undergoing digital transformation. In this paper, we evaluate a variety of commercially available, lightweight large language models (LLMs) for sentiment, emotion, aspect-based sentiment, and toxic speech detection on noisy, real-world social media datasets. These lightweight models are compared (against traditional RoBERTa-based classifiers and heavyweight “pro” LLMs) in terms of performance, latency, and cost trade-offs that are key for scalable deployment across the enterprise. Our results show that lightweight LLMs achieve good accuracy with considerably lower response times and costs, which makes them suitable as the next step in digital transformation for real-time social media analytics. Additionally, we investigate the importance of prompt engineering and show that preprocessing can be very limited for LLMs. This study provides operational guidelines regarding model selection based on latency, accuracy, and cost trade-offs as well as optimal prompt engineering techniques. Finally, it evaluates the intersection of operational performance and resource efficiency to help researchers and developers integrate LLM-based social media analytics into digitized business practices.

Ioannis Kapantaidakis, Ioannis Kopanakis, Emmanouil Perakakis et al. · 0 citations
Open access Jul 2026

Comparison of Machine Learning Algorithms for Sentiment Analysis of Trans Jogja on Social Media

The social media platform X has become an important channel for the public to express opinions and share experiences regarding public services, including Trans Jogja. User-generated content from this platform provides valuable insights into public perceptions of service quality. However, because these data consist of unstructured text, sentiment classification techniques based on machine learning are required to analyze them effectively. This study aims to compare the performance of several machine learning algorithms for sentiment classification, including Naïve Bayes, Support Vector Machine (SVM), Random Forest, Neural Network, Logistic Regression, and Decision Tree, in classifying user sentiment toward Trans Jogja on the X platform. Data were collected through a web crawling process using Tweet Harvest with keywords related to Trans Jogja, covering the period from January 1, 2025, to June 10, 2026, resulting in a dataset of 3,035 tweets. The preprocessing stage included data cleaning, case folding, tokenization, normalization, stopword removal, and stemming. Text representation was performed using the Term Frequency–Inverse Document Frequency (TF-IDF) method. The dataset was then divided into training and testing sets using five train–test split ratios: 90:10, 85:15, 80:20, 75:25, and 70:30. Model performance was evaluated using a confusion matrix and the corresponding accuracy, precision, recall, and F1-score metrics. The experimental results demonstrate that the Support Vector Machine (SVM) consistently outperformed the other algorithms across different data split ratios. At the 85:15 train–test split, the SVM achieved an accuracy of 91%, precision of 91%, recall of 91%, and an F1-score of 91%, indicating that it is the most effective algorithm for sentiment classification of Trans Jogja users on the X platform.

Putri Muryanti Setyowati, Y. Pristyanto, Arif Nur Rohman · 0 citations
Preprint Aug 2026

Multiclass Sentiment Analysis for Identifying Political Viewpoints

The rapid growth of social media has created vast amounts of political discourse, which provides valuable opportunities to analyze public opinions and identify different political perspectives. Sentiment Analysis (SA) is a core task in Natural Language Processing (NLP) that allows the computational study of attitudes and opinions in textual data, and has become increasingly important for understanding political discourse. In this work, we investigate multiclass sentiment analysis of political view- points on social media, that is to automatically discriminate multiple sentiment classes over political issues and figures. To solve this task we design and evaluate two machine-learning approaches based on XGBoost and BERT. We train and evaluate the models on a labeled dataset of political social media posts using standard classification metrics. The experimental results show that the XGBoost model reaches an F1-score of 0.2835 and the BERT- based model reaches an F1-score of 0.2806 on the test set. These results demonstrate the challenge of classifying complex and contextualized political discourse sentiment and provide a baseline for future research in multiclass political sentiment analysis.

G. Bade, O. Kolesnikova, J. Oropeza et al. · 0 citations
2026

Applying machine learning to social media analysis: opportunities and challenges

This article applies a machine learning approach to analyzing social media, which has a decisive impact on user behavior, public sentiment, and current trends. The key pillars of the approach include natural language analysis, text classification, sentiment analysis, and the identification of hidden patterns. Particular attention is given to the practical applications of such technologies in marketing, political analysis, public opinion monitoring, and reputation management. Along with the benefits, key challenges are discussed, such as data quality issues, privacy concerns, ethical limitations, and model interpretability. A conclusion is reached regarding the need for a comprehensive approach to implementing machine learning in social media analysis, taking into account technological and social factors.

Raisa S.-A. Gatsayeva, Z. Batchaeva, M. G. Plotnikova · 0 citations
Open access Jul 2026

Machine learning-based sentiment analysis to determine social perception of Generation Z on Twitter/X

The arrival of Generation Z in the social and work environment has sparked a heated debate on social media, marked by a division between innovative visions and critical stereotypes. This research develops a sentiment analysis model based on machine learning to understand public perception of this population group. Using a dataset obtained from Twitter/X through content and data extraction from the web with the Octoparse tool, three classification algorithms were trained and evaluated: Support Vector Machines (SVM), Random Forest, and Long Short-Term Memory (LSTM). Due to the class imbalance inherent in generational discussions, data balancing techniques (SMOTE, ADASYN) were applied. The results indicate that the Optuna-optimized LSTM model with SMOTE+Tomek balancing performed best with an accuracy of 82.39%, outperforming classical approaches. This tool allows for the identification of opinion trends (positive, negative, or neutral) about Generation Z, providing valuable input for sociologists and organizations seeking to understand the dynamics of “Generation Z” or “Centennials.”

Hugo Vega-Huerta, Luis Casaperalta-Pacheco, Frida López-Córdova et al. · 0 citations
Open access 2026

Technical Research on Social Media Text Stance Detection

This paper will explore the construction of stance data, the evolution of model paradigms, and the issue of generalization in real-world applications, and propose new theoretical paths to improve the practicality and credibility of this field.

Jiayi Zhao · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.