Skip to content

Similar papers

Review Open access Jul 2026

Detecting Hate Speech in Hindi Digital Discourse Using Transformer–Long Short-Term Memory Models

Hate speech on digital communication platforms has become a major obstacle to healthy online discourse, especially in multilingual societies such as India, where Hindi is a dominant language in social media interactions. Hostile, offensive, and defamatory speech is linguistically and socio-culturally complex because of colloquial idioms, regional variations, code-mixing, and culturally embedded references. However, effective detection is crucial for creating safer and more inclusive digital communication environments. This study evaluates advanced language models for analysing hate speech in Hindi social media content. A dataset of 21000 Hindi posts from Twitter and public repositories, categorized into five categories, was analysed: hate, offensive, fake, defamation, and non-hostile. General-purpose models were tested against Hindi-specific language models, including Hindi-BERT (based on Bidirectional Encoder Representations from Transformers, or BERT) and MuRIL, to investigate whether performance can be further enhanced by integrating Long Short-Term Memory (LSTM) layers. In tests, the best-performing model, Hindi-BERT (with sequential learning added), correctly identified 94 out of every 100 posts—a considerable improvement over simpler models. For social media platforms, it has real-world consequences: it can automatically identify harmful content to be reviewed by a human, minimize nuisance alerts that waste human moderators' time, and identify defamatory or bogus posts before they gain a lot of traction. The results offer a systematic approach to researching how hostility develops in the Hindi-speaking online community, how linguistic creativity (e.g., slang, sarcasm, code-mixing) can conceal or manifest hostility, and how decisions about content moderation influence public discourse for communication scholars. Overall, this paper illustrates that language-specific computational tools can be used for both platform governance and communication research, provided that the cultural context is considered. Finally, technical methods are combined with communication scholarship to explain and curb harmful speech in the online public sphere.

Rachna Narula, Vedika Gupta, Jawad Khan et al. · 0 citations
Review Open access 2026

Culturally Aware Malay–English Code-Mixed Hate Speech Detection: A Systematic Review and Research Taxonomy

Social media platforms such as X, formerly Twitter, have become major channels for communication, information sharing and public discussion. However, the rapid growth of user-generated content has also increased the spread of hate speech and offensive language. Automated hate speech detection remains challenging in multilingual and code-mixed environments, where users frequently combine languages, informal spelling, slang, abbreviations and culturally specific expressions. In Malaysia, online discourse often involves Malay-English code-mixing, commonly referred to as Manglish, which creates additional challenges for natural language processing systems. This study presents an evidence-informed systematic review and research-readiness taxonomy for culturally aware Malay-English hate speech detection. Unlike conventional reviews that mainly summarize model architecture and performance scores, this review critically evaluates existing studies based on dataset availability, code-mix authenticity, annotation practice, cultural sensitivity, model architecture, evaluation strategy, explainability, robustness and deployment readiness. To strengthen this review, this article incorporates a completed empirical case study on Manglish hate-speech detection, using posts collected from X (formerly Twitter). The case study used keyword-based data collection, Malaya NLP-based language filtering, bilingual manual annotation, manual class balancing and transformer-based evaluation using mBERT, XLNet and XLM-RoBERTa. The case evidence is used only as an empirical lens to illustrate practical challenges in dataset curation, class imbalance, lexical overlap and model robustness. It is not positioned as a new benchmark dataset or an independent experimental contribution. The review finds that existing Malay and Malay-English resources remain fragmented. Some datasets are monolingual Malay hate speech datasets, some are bilingual but language-separated, while others are code-mixed but developed for sentiment analysis rather than hate speech detection. Transformer-based models such as BERT, mBERT and XLM-RoBERTa show strong potential, but their results are difficult to compare due to inconsistent datasets, label definitions, class distributions, evaluation metrics and limited robustness testing. The main novelty of this review is the proposed taxon- omy that evaluates the field through six readiness dimensions, which are data, linguistic, cultural, modelling, evaluation and deployment readiness. The findings from this study provide a structured foundation for developing culturally aware, robust, explainable and deployable hate speech detection systems for Malay-English code-mixed social media.

F. Azmi, Normaisharah Mamat, Rawad Abdulghafor et al. · 0 citations
Open access Jul 2026

Low-Resource Hate Speech Detection in English-Swahili Code-Switched Text Using Fine-Tuning of Pre-trained Language Models

This study explores a low-resource approach to detecting hate speech in English and Swahili code-switched text by fine-tuning pre-trained language models, and shows that fine-tuning modern language models can offer a practical and scalable solution for hate speech detection in multilingual environments.

Kipkebut Andrew, Jepkemei Betty · 0 citations
Open access Aug 2026

Translanguaging and identity construction among Indonesian–English bilinguals on X social media platform: A computer-mediated discourse analysis

Background: The evolution of communication in the digital era has led to increasingly flexible and dynamic language practices in everyday interactions. Social media, serving as a virtual public space, enables bilingual and multilingual users to combine multiple languages within a single utterance, a phenomenon known as translanguaging. Aims: This study aims to identify the forms of translanguaging in posts on the X platform and analyze their role in shaping users’ digital identities during institutional crises. Methods: The research employs a qualitative approach using Susan Herring’s (2004) Computer-Mediated Discourse Analysis (CMDA) framework. The dataset consists of a small purposive corpus of seven posts discussing academic sexual harassment cases on Indonesian X, collected through documentation and analyzed through coding, categorization, and interpretative analysis. Results: The findings revealed that code-mixing is the dominant form of translanguaging, accompanied by limited instances of hybrid language forms. These linguistic practices function to express emotions, reinforce opinions, deliver social criticism, and address systemic institutional issues. The examined cases indicate that translanguaging serves as a highly deliberate, context-sensitive discursive strategy to navigate context collapse, manage audience design, and establish positioning in digital interactions. Furthermore, language choices project fluid, multifaceted personas, shifting between intellectual, critical, and empathetic positionings. Implications: The study demonstrates that translanguaging is not merely a linguistic phenomenon but also a social practice closely tied to citizen agency, self-representation, and the dynamics of digital communication during crisis situations.

M. Abdurrahman · 0 citations
Review Open access Aug 2026

Artificial Minds, Cultural Shadows: Cultural Alignment, Identity, and Voice Across Multiple Large Language Models

Comparison of five widely used large language models suggests that AI-generated language may shape how culturally situated perspectives are expressed, with differences across models indicating that AI-generated language may shape how culturally situated perspectives are expressed.

Ashkan Goudarzi, Aylar Naderi Zonouz · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.