Jul 2026· Journal for Language Technology and Computational Linguistics· Vol 39, pp. 107-152· 0 citations
Abstract
Stance classification in Natural Language Processing (NLP) is not just an academic exercise but a crucial tool for understanding political discourse and the attitudes underlying political statements. This research addresses the challenge of limited annotated datasets in political science by proposing a practical sentence-level dataset sourced from professional politicians for binary subjective stance classification - support or oppose - using bootstrapping in a SetFit model. The study leverages the Sentence Transformers architecture and incorporates traditional linguistic approaches to enhance explainability. We employ corpus linguistics, tailored lexicons, and lexicogrammatical rules to identify key linguistic features such as positive affect, negative affect, pro polarity, con polarity, certainty, emphatics, doubt, hedges. SHAP analysis quantifies the influence of these features on SetFit model decisions. Our findings demonstrate that iterative bootstrapping significantly enhances the efficacy of few-shot learning in subjective stance classification, and we highlight the importance of linguistic features, particularly pro/con polarity and affective expressions. The StanceSentences dataset and our hybrid analytical approach offer a benchmark for future research, emphasizing the need for nuanced, multi-layered analysis in political discourse.
The rapid growth of social media has created vast amounts of political discourse, which provides valuable opportunities to analyze public opinions and identify different political perspectives. Sentiment Analysis (SA) is a core task in Natural Language Processing (NLP) that allows the computational study of attitudes and opinions in textual data, and has become increasingly important for understanding political discourse. In this work, we investigate multiclass sentiment analysis of political view- points on social media, that is to automatically discriminate multiple sentiment classes over political issues and figures. To solve this task we design and evaluate two machine-learning approaches based on XGBoost and BERT. We train and evaluate the models on a labeled dataset of political social media posts using standard classification metrics. The experimental results show that the XGBoost model reaches an F1-score of 0.2835 and the BERT- based model reaches an F1-score of 0.2806 on the test set. These results demonstrate the challenge of classifying complex and contextualized political discourse sentiment and provide a baseline for future research in multiclass political sentiment analysis.
G. Bade, O. Kolesnikova, J. Oropeza et al.· 0 citations
This paper investigates the affect heuristic in online argumentative discourse through a corpus-based study combining manual and automatic annotation. Drawing on research in argumentation theory, psychology, and decision science, the study approaches affect not as an external addition to reasoning but as a recurring component of evaluative judgment. The analysis focuses on discussions of climate change and artificial intelligence collected from Reddit and X (formerly Twitter), domains characterised by uncertainty, risk perception, and public controversy. The study employs a bottom-up annotation methodology in which four human annotators identify instances of affect heuristic and related cognitive biases in a corpus of more than 30,000 posts and comments. Inter-annotator agreement is assessed using Fleiss’ κ, Cohen’s κ, Gwet’s AC1, and weighted F1 measures. In a second stage, the same annotation scheme is applied to GPT-4o, treated as a constrained fifth annotator operating within predefined categories and probabilistic classification rules. The results show that the affect heuristic is the most frequent heuristic pattern in the corpus, occurring more often than confirmation bias, availability heuristic, or representativeness heuristic. Although traditional κ coefficients remain low because of category imbalance, agreement measures robust to prevalence effects indicate substantial consistency among annotators. The automatic annotation stage reveals partial alignment between human and model judgments, while also exposing systematic discrepancies in the model’s distribution of categories. A lexical and discursive analysis further demonstrates that affect heuristics do not necessarily manifest through explicit emotion vocabulary. Instead, they frequently appear through evaluative framing, practical reasoning under uncertainty, and subtle stance-taking related to collective action and future-oriented judgment. The findings contribute to empirical research on emotional processes in argumentation and demonstrate how affective reasoning can be operationalised and studied through combined qualitative and computational methods.
Paulina Żelewska, Barbara Konat· Człowiek i społeczeństwo· 0 citations
Sarcasm detection is a complex task in natural language processing because it depends on implicit mood variations, contextual comparison, and the congruence of polarity between the literal form of expression and its intended meaning. Although the transformer-based models, including BERT and DeBERTa, have been taking a great leap in performance by introducing contextual self-attention, they mainly learn sarcasm patterns with an implicit hypothesis of sentiment polarity contradictions, which define the discourse of sarcasm. In this paper, a polarity-sensitive transformer model that explicitly incorporates sentiment data in representation learning is presented to detect sarcasm. In contrast to traditional fine-tuning methods, which treat sarcasm as a generic classification problem, the methodology adds sentiment-polarity cues to the embedding space, enabling the model to fine-tune contextual representations in a polarity-sensitive way. The proposed method will improve the ability of semantic representations of a context to reflect incongruity patterns in contextual segments. The positive results of experiments on benchmark sarcasm datasets indicate that explicit polarity integration is more robust and generalizes better than traditional transformer baselines, particularly in context-specific situations. The findings indicate that embedding-level sentiment improvement offers a sound theoretical and practical guideline in the process of expanding sarcasm detection beyond implicit contextual modeling.
S. Nagini, Karnam Akhil, Harshitha Upadhyayula et al.· International journal of com...· 0 citations
Although the impact of digital news media on public opinion has grown significantly, there are currently no computational frameworks that empirically connect news mood to election results. This research suggests an artificial intelligence (AI)-based solution that uses deep learning and natural language processing to analyze how political news sentiment affects election outcomes. A dataset of 17,000 political news stories gathered during the general elections in Pakistan in 2024 was manually classified into three categories: good, negative, and neutral. In order to capture contextual and temporal relationships in political speech, the texts underwent normalization, tokenization, lemmatization, and modeling using a Bidirectional Long Short-Term Memory (BiLSTM) network. To evaluate media influence, party-level mood indices were calculated using model predictions and statistically associated with official vote shares. According to experimental results, the suggested BiLSTM model outperforms conventional machine learning baselines and reaches an accuracy of 89%. Additionally, sentiment indices and electoral results show a statistically significant link, suggesting a quantifiable relationship between media mood and voter behavior. The suggested paradigm offers empirical insights into the function of news sentiment in election dynamics and a repeatable technique for extensive political media study.
Yawar Abbas Abid, Muhammad Kashif, Javed Ferzund et al.· PeerJ Computer Science· 0 citations
As the ecosystem encompassing social media and product reviews grows ever more intricate, emotional expression presents prominent traits including subtlety, sarcasm, fragmentation and multimodal fusion. Traditional machine learning models (e.g., SVM), which rely on manual feature engineering, encounter bottlenecks in recognition accuracy when dealing with irony, metaphor, and long-distance sentiment dependencies. Sentiment analysis of reviews is thus trapped in the dual predicament of "semantic noise" and "shallow understanding." This paper focuses on the advantages and accuracy verification of Large Language Models (LLMs) in tackling high-difficulty sentiment analysis of reviews. This study abandons the mere enumeration of single accuracy values and instead delves into the cognitive breakthroughs of LLMs across three key dimensions: context-aware ambiguous meaning resolution, which deeply interprets ironic and euphemistic connotations; fine-grained sentiment element extraction, accurately identifying the polarity of praise or criticism toward specific product attributes; and closed-loop verification through sentiment generation and explanation, providing traceable justifications via chain-of-thought mechanisms.
Yaxuan Wang· Applied and Computational En...· 0 citations
It is demonstrated that lemmatization does not produce uniform gains across architectures: while linear models and croBERT display small but measurable improvements from morphological normalization, non-linear models such as RBF SVM and neural networks experience substantial declines in performance.
I. Ljubi, M. Horvat, G. Gledec et al.· Electronics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.