Skip to content
Conference

IG-Guided Synonym Substitution Attacks on IndoBERT AI Detection Across Domains

Jul 2026 · 2026 IEEE International Conference on Industry 4.0, Artificial Intelligence, and Communications Technology (IAICT) · pp. 664-669 · 0 citations · 29 references

Abstract

The rapid adoption of large language models has urged the development of reliable detectors that are capable of distinguishing AI-generated text from human-written content. While recent Transformer-based detectors have shown promising performance, their robustness against adversarial manipulation remains underexplored, particularly in multilingual and cross-domain settings. This study investigates the vulnerability of an IndoBERT-based AI-generated text detector to synonym substitution attacks guided by Integrated Gradients (IG) that identify words most influential to the model’s predictions. By leveraging IG to selectively perturb high-importance tokens, we construct a constrained synonym substitution attack that aims to evade detection while preserving semantic fidelity. Experiments are conducted on Indonesian news articles and speech transcripts to assess domain-specific robustness. The results reveal that attribution-guided attacks can significantly degrade detector performance, achieving meaningful attack success rates on AI-generated texts that were initially classified correctly. Moreover, the noticeable cross-domain behaviors are also observed. Where speech texts are more vulnerable to meaning-preserving perturbations but require substantially higher attack effort, whereas news texts demand fewer attempts at the cost of higher lexical modification. Overall, this work proves that strong pre-attack accuracy does not guarantee the model’s resilience against guided adversarial attacks and emphasizes the importance of incorporating explainability-driven adversarial analysis in the development of future detection systems.

View source