Skip to content
Book Open access

OTPCL: Optimal Transport Driven Pseudo-Labeling with Contrastive Learning for Social Bot Detection

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 5821-5832 · 0 citations · 24 references

Abstract

Social bot detection is vital for protecting online platforms from misinformation and manipulation. In recent research, graph neural networks (GNNs) have emerged as a powerful approach, since they leverage relational patterns and social interactions to identify coordinated bot behaviors. However, two key challenges arise from the nature of real-world social networks: First, bots often actively interact with human users through follows, replies, and mentions, creating numerous ''heterophilous'' edges, i.e., connections between different classes. These cross-class ties disrupt the homophily assumption underlying many GNNs, causing messagepassing to propagate and amplify errors. Second, due to the high cost and time required for manual annotation, social media platforms typically contain a large proportion of unlabeled data, with only a small fraction labeled for bot detection. Unlabeled data are often underutilized, making supervision sparse. To address this, we propose OTPCL (Optimal Transport Driven Pseudo-Labeling with Contrastive Learning), a plug-in framework for GNN-based social bot detection. OTPCL first employs contrastive learning to obtain well-separated node representations. It then formulates pseudolabel assignment as an optimal transport problem, which simultaneously generates pseudo-labels and quantifies their reliability via transport scores. These scores guide two key mechanisms: selective removal of unreliable heterophilous edges to purify the graph structure, and reducing the influence of pseudo-labels with transport scores below the dynamic threshold. Extensive experiments on three widely used benchmark datasets demonstrate that OTPCL consistently improves the detection performance across six different GNN backbones, showing strong robustness and generalization in both low-labeled and fully-labeled scenarios.

Read PDF

Similar papers

Book Open access Aug 2026

Rethinking Generalization in Graphs: A Hierarchical Interaction Perspective for Generalist Detection

With the increasing heterogeneity of social networks and online interaction systems, generalist graph anomaly detection (GAD) has become essential for identifying abnormal and fraudulent behaviors in complex environments. However, most existing GAD approaches rely heavily on domain-specific semantic alignment, which substantially restricts their ability to learn transferable node representations and often leads to poor generalization on unseen graph domains. To address this challenge, we propose HIerarchical Interaction MOdeling for zero-shot generalist GAD (termed HIMO-GAD). HIMO-GAD enables anomaly detection across diverse graph domains without retraining or access to target-domain supervision by modeling the evolutionary trajectories of node representations across hierarchical structural depths, thereby capturing interaction patterns that exhibit strong cross-domain stability. Specifically, HIMO-GAD integrates two core components: (1) a Dynamic Interaction Modeling Module that characterizes cross-layer interaction evolution to extract transferable representations, and (2) an Anomaly-Aware Regulation Mechanism that combines gradient immunity and centralization regularization to suppress overfitting and stabilize cross-domain generalization. Extensive experiments on multiple real-world graph datasets demonstrate that HIMO-GAD consistently outperforms state-of-the-art baselines in strict zero-shot settings, achieving up to a 10% improvement in key evaluation metrics and exhibiting strong generalization across heterogeneous graph domains.

Xiangping Zheng, Xuan Feng, Bo Wu et al. · 0 citations
Open access 2026

Sparse Structural Knowledge Enhanced Graph Neural Networks for Anomaly Detection in Social Networks

: Social network platforms have become primary channels for information dissemination, yet they are increasingly exploited by anomalous users such as bots, fake accounts, and coordinated disinformation spreaders. These malicious actors manipulate public opinion, spread misinformation and undermine platform integrity, posing severe threats to the security of the online ecosystem. Accurate detection of such users is challenging because they often organize into sophisticated high-order connection patterns that extend beyond local neighborhoods. Existing methods address this by either injecting predefined motifs as handcrafted features, which lack flexibility to discover unknown patterns, or employing higher-order Graph neural networks (GNNs) at prohibitive costs. Crucially, neither method treats structural information as learnable knowledge that can be automatically acquired from data and explicitly represented. To bridge this gap, we propose SparseGNN, a structural-knowledge-enhanced framework for anomalous user detection. It regards atomic subgraph patterns as fundamental, learnable units of structural knowledge. This framework is concatenated with original node features and fed into any standard GNN, without modifying the backbone architecture. Experiments on real-world datasets demonstrate that SparseGNN improves the accuracy and F1-score of standard GNNs for anomalous users detection without requiring predefined patterns, while maintaining linear complexity. Because the learned atomic patterns capture global high-order topology, the resulting structural knowledge representation is inherently less sensitive to localized edge perturbations, incidentally conferring improved stability under adversarial structural attacks.

Zehan Li, Yingyi Li, Zhiwei Tang et al. · 0 citations
Open access Jul 2026

A Heterogeneous Graph Attention Network with Pre-trained Language Models for Multi-Modal Fake News Detection

Heterogeneous Graph Attention Network (HGAT) is proposed, where a pretrained BERT-Large encoder is coupled with a Heterogeneous Graph Attention Network (HGAT) to learn joint representations for textual, social network and external knowledge graph features.

Akash Garg, Sachin Pachauri · 0 citations
Open access Jul 2026

FakeDiverse a curated multi-source news corpus for context-aware fake news detection using BERT and DeBERTa

This study examines the effectiveness of two transformer-based architectures—BERT and DeBERTa—for identifying fake news using only textual information from headlines and article bodies and achieves strong performance on FakeDiverse corpus, demonstrating the need for enhanced generalization strategies as well as domain adaptation.

A. Kumar, A. S, Akshara G. Bhat et al. · 0 citations
Open access Aug 2026

4HAN: An Enhanced Neural Network for Fake News Detection using Hypergraph

A Four-Level Hierarchical Attention Network (4HAN) that incorporates word-, sentence-, and headline-level attention, along with Hypergraph Convolution and Hypergraph Attention, is proposed using the LIAR dataset and shows a detection accuracy rate of 96.00%, which beats multiple existing methodologies in fake news detection.

Alpana A. Borse, Gajanan K. Kharate, N. Wasatkar · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.