Aug 2026· Expert systems· Vol 43· 0 citations· 18 references
TL;DR
A novel abstractive summarisation framework designed to distill coherent and semantically rich summaries from social media discussions that consistently outperforms mainstream summarisation methods, including advanced systems like ChatGPT, particularly in preserving semantic alignment and improving readability under noisy conditions.
Abstract
The exponential growth of social media has reshaped global communication and decision‐making in business, politics and economics. Yet, the sheer volume and informal, unstructured nature of user‐generated content present major challenges for meaningful analysis. This study introduces a novel abstractive summarisation framework designed to distill coherent and semantically rich summaries from social media discussions. Built on the T5 transformer architecture and enhanced through targeted transfer learning, the system effectively captures the fragmented, slang‐rich language patterns common across platforms. Evaluation is conducted using a suite of semantic‐aware metrics—including ROUGE‐WE, SUPERT and Shannon entropy—alongside human‐centric criteria such as coherence, fluency, consistency and lexical diversity. Results show that the proposed model consistently outperforms mainstream summarisation methods, including advanced systems like ChatGPT, particularly in preserving semantic alignment and improving readability under noisy conditions. Comparative analysis underscores the framework's robustness in handling unstructured, domain‐specific discourse. These findings position the model as a valuable tool for real‐time, high‐volume social media analytics. Future work will explore hybrid neural quality assessment and interactive feedback mechanisms to further enhance domain adaptability and summary fluency.
This paper presents a framework that preserves semantics in LLM-based opinion summarization while minimizing token usage and computational cost and demonstrates that this method significantly reduces token usage and computational cost while consistently outperforming traditional AI-based and standard LLM summarization baselines in terms of content coverage, balance, and semantic preservation.
This survey presents a systematic review of 121 references spanning 2002 to 2026, tracing the evolution of TextRank-based approaches into hybrid LLM pipelines and advancing three qualified arguments.
Ahmed J. Jabur, Asmaa Abdul Azeez Dakhil, Israa Saad Mohammed et al.· Iraqi Journal for Computers...· 0 citations
A growing convergence between social science and machine learning enables, in principle, large-scale analyses of complex social phenomena through text. Yet, approaches leveraging supervised text classification based on human-annotated data for statistical analysis often treat conceptual validity and technical performance as separate challenges, impairing measurement quality. We provide guidelines to bridge this gap in what we call computational social mixed methods pipelines across three stages: data annotation, model training, and statistical analysis. Building on best practices and our own methodological innovations, such as “Iterative Annotation” and “Training on Confident Examples”, we address recurring pitfalls like unbalanced training data or stagnant model performance. We also discuss when large language models constitute a viable alternative to transformer-based classifiers. Using a case study on countering online hate, we illustrate how consequently integrating social science and machine learning expertise improves the validity and comparability of computational social science.
Alina Herderich, Jana Lasser, Mirta Galesic et al.· Behavior Research Methods· 1 citation
. Social media sentiment analysis faces a persistent aggregation problem: lexicon-based and transformer-based models often produce inconsistent outputs for the same short, informal, and stylistically heterogeneous texts. This paper introduces ADRTW (Adaptive Dynamic Reliability-Trig-gered Weighting), an interpretable sentiment fusion framework that combines heterogeneous sentiment estimators using rule-guided reliability weights derived from textual cues, inter-model disagreement, and consistency patterns [5, 8]. The framework is evaluated on a Reddit dataset containing 1,577 posts, 354,050 comments, and 187,666 authors collected between 2017 and 2025, together with a controlled synthetic benchmark for aggregation comparison. The results show that ADRTW remains competitive with static averaging in controlled settings while preserving context-sensitive local variation in large-scale discourse analysis. Beyond sentiment fusion, the ADRTW-derived signal supports complementary analyses of online discussions, including temporal trend inspection, toxicity-aware interpretation, and participation-based clustering. Overall, the proposed framework provides a transparent and reusable basis for examining emotional dynamics in social media discourse.
Aniko Apro, L. Sasi· Annales Mathematicae et Info...· 0 citations
Climate change news coverage varies by region and media organization. These differences show in the way news is framed, the themes chosen, and communication priorities. This study suggests a scalable and unsupervised natural language processing (NLP) framework to measure sentiment changes and framing patterns in global climate discussions across major news outlets. The proposed system uses Zero-Shot BART-large-MNLI as the main sentiment classification model, while VADER serves as a comparative baseline during model validation. Key themes are identified using BERTopic. We detect changes in climate narratives over time with the Pruned Exact Linear Time (PELT) changepoint detection algorithm. We introduce a dual-baseline framework to calculate relative sentiment changes. This combines regional consensus with a scientific baseline. It allows comparison of framing patterns without hiding systematic regional differences. The BART-MNLI model was tested against a manually annotated sample of articles. It achieved an accuracy of 98.0%, an F1-score of 0.98, and Cohen’s
κ
of 0.96, showing strong agreement with human annotations. It significantly outperformed lexicon-based sentiment analysis for professional news text. We also assessed statistical robustness with the Kruskal-Wallis H-test, bootstrap confidence intervals, and PELT sensitivity analysis. The framework reveals regional sentiment differences, outlet-level framing patterns, thematic trends, and notable shifts in response to major climate events. It offers a clear, reproducible, and flexible method for analyzing global climate communication.
Unknown authors· Frontiers in Climate· 0 citations
This study addresses the challenge of achieving consistent and scalable quantification of semantic divergence in large-scale heterogeneous textual data by proposing an integrated solution that combines methodological design and system architecture. Existing approaches primarily rely on lexical statistics or vector-based distances, which are insufficient for capturing the global structure of high-dimensional semantic spaces and lack consistency across different text granularities. To overcome these limitations, we propose the Core Semantic Variance Index (CSVI), which leverages semantic embeddings together with Principal Component Analysis (PCA) and the Participation Ratio (PR) to characterize the distribution of semantic variance. In addition, a distributed semantic analysis service platform is developed to support unified analysis and efficient computation across diverse text structures. Comparative experiments against traditional lexical-based and existing embedding-based methods show that CSVI achieves accuracies of 86.35%, 93.05%, and 88.32% on Multilingual-STSB, Python Dialogue, and 20 Newsgroups, respectively, demonstrating its effectiveness and applicability within distributed service environments.
Yen-Yu Chang, Meng-Che Tsai, Yu-Chen Chien et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.