It is demonstrated that clean-text performance is not a reliable predictor of adversarial robustness, and the results underscore the necessity for architecture-specific defences and frame smishing detection as an adversarial cybersecurity challenge rather than a static classification task.
Abstract
Smishing detection systems are commonly trained and evaluated on clean, monolingual text. In low-resource settings, however, attackers frequently circumvent these systems through character obfuscation, cross-lingual code-switching, and structural perturbation. This study evaluates adversarial robustness for five model architectures: three classical lexical models (Random Forest, XGBoost, CNN+BiLSTM) and two multilingual transformers (mBERT, XLM-RoBERTa), using a dataset of 27,037 messages. Classical models are subjected to black-box generic attacks, while transformers are evaluated with attention-guided targeting. Each model is tested across three attack types and intensity levels, with performance measured by the Robustness Degradation Ratio (RDR). The results reveal a distinct architectural boundary: classical models experience near-catastrophic failure under character obfuscation and structural perturbation (RDR up to 0.988), whereas transformers demonstrate significantly greater resilience (RDR up to 0.351), with structural perturbation representing their most pronounced vulnerability. Effect-size analysis (Cliff's d) indicates a substantial difference between the two model categories. Within the transformer group, XLM-RoBERTa, despite achieving a higher clean-text baseline, exhibits greater degradation than mBERT. These findings demonstrate that clean-text performance is not a reliable predictor of adversarial robustness. Statistical validation using Mann-Whitney U and Friedman tests confirms that these patterns are attributable to model architecture rather than sampling. The results underscore the necessity for architecture-specific defences and frame smishing detection as an adversarial cybersecurity challenge rather than a static classification task.
Dual Modular Redundancy (DMR) and Triple Modular Redundancy (TMR) are commonly used methods for providing fault detection and/or tolerance in safety-critical systems by incorporating redundant – and often diverse – components. However, these systems can still be susceptible to adversarial attacks that may deceive AI models, potentially leading to severe consequences. In this paper, we introduce enhanced DMR and TMR strategies for image-based object detection, leveraging image transformations during inference to help reduce the impact of adversarial inputs, while preserving the inherent advantages of diverse redundancy for safety purposes. Experimental results demonstrate that our approach significantly improves robustness under adversarial conditions, achieving up to 12.9% and 12.2% higher accuracy than state-of-the-art solutions in DMR and TMR configurations, respectively, when attacks are individually crafted for each image. Furthermore, against universal adversarial attacks, our solution achieves even greater accuracy gains, with up to 26.8% and 26.0% higher accuracy in DMR and TMR configurations, respectively.
Martí Caro, Axel Brando, Jaume Abella· ACM Transactions on Design A...· 0 citations
This paper proposes Homoglyph-Guided Beam Search (HG-BS), an adversarial attack framework that generates evasive URLs preserving both visual appearance and functional validity under strict structural constraints, and establishes that current high-accuracy URL detectors rely on fragile token patterns rather than robust semantic understanding.
Junhyeong Lee, Hyun Kwon· International Journal of Mac...· 0 citations
This work presents a two-phase evaluation of ten Llama variants using the OWASP Top 10 for LLM Applications, and applies nine encoding obfuscations to the same prompts, which fully bypasses all text-only models.
Nourin Shahin, Izzat Alsmadi· Practice and Experience in A...· 0 citations
The Adversarial-Resilient Lightweight Random Forest (AR-LRF) model is proposed, combining controlled ensemble complexity with simulated adversarial perturbations applied during training to mitigate adversarial vulnerabilities.
A. Chaudhuri, M. B· Scientific Reports· 0 citations
This paper introduces the first attack that directly optimizes an encoder-attention objective under an imperceptible, bounded, bounded perturbation, and argues that encoder attention concentrates the model's spatial reasoning, so corrupting it propagates through the detection pipeline more disruptively than perturbing the detection output alone.
Ridma Jayasundara, Shaheer Mohamed, Tharindu Fernando et al.· 0 citations
This paper evaluates the robustness of the Support Vector Machine (SVM) classifier, a leading algorithm in state-of-the-art HT detection frameworks, under gradient-based adversarial attacks, and highlights the need to reframe hardware security evaluations beyond nominal accuracy toward adversarial robustness.
Ashutosh Ghimire, Lingwei Chen, Cole Castronova et al.· Journal of electronic testin...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.