DAMSE: a dialect-aware multi-strategy ensemble framework for Arabic vishing detection with zero-shot learning
TL;DR
DAMSE is validated on two linguistically distinct datasets and establishes the first zero-shot baselines for vishing detection, introducing dialect-adaptive ensemble weighting that provides a consistent gain over uniform fusion, and releasing the first comprehensive multi-dialect Arabic vishing dataset with structured risk indicator annotations.
Abstract
Voice phishing (vishing) attacks targeting Arabic and Korean speakers represent a growing cybersecurity threat that demands automated detection systems capable of operating across diverse linguistic contexts. This article proposes Dialect-Aware Multi-Strategy Ensemble (DAMSE), a novel framework that synergistically fuses five complementary classification strategies for multilingual vishing detection: fine-tuned transformer models (Arabic BERT/Korean BERT), enhanced zero-shot classification with multi-template hypothesis averaging and Platt scaling calibration, few-shot learning via SetFit contrastive training, risk-score gradient boosting incorporating domain-specific social engineering indicators, and dialect-adaptive weighted ensemble fusion optimized per regional variant. We validate DAMSE on two linguistically distinct datasets: a curated Arabic vishing corpus comprising 448 multi-turn conversations across nine dialects (Modern Standard Arabic (MSA), Egyptian, Gulf, Jordanian, Saudi, Yemeni, Sudanese, Iraqi, Syrian) and 23 fine-grained categories, and the Korean KorCCVi v2 dataset containing 23,550 real-world call transcripts with severe class imbalance (1:13.3 ratio). Under stratified 5-fold cross-validation with repeated runs (five seeds × five folds = 25 runs), DAMSE achieves 99.31 ± 0.53% accuracy on Arabic and 99.92 ± 0.05% on Korean, demonstrating robust cross-lingual generalizability with tight variance bounds. On the held-out test splits, DAMSE achieves 99.99% accuracy on Arabic and 99.80% on Korean with zero false positives on both datasets. Key contributions include: establishing the first zero-shot baselines for vishing detection (92.00% Arabic, 94.20% Korean), achieving 96.50% accuracy with only five examples per class through data-efficient few-shot learning, introducing dialect-adaptive ensemble weighting that provides a consistent gain over uniform fusion, and releasing the first comprehensive multi-dialect Arabic vishing dataset with structured risk indicator annotations. Explainable artificial intelligence (AI) analysis reveals convergent vishing patterns across languages: financial terminology dominance, urgency markers, formal impersonation language, and extended narrative length. Comparison with 15 state-of-the-art approaches confirms DAMSE’s advancement across performance, zero-shot capability, multi-dialect support, and cross-lingual generalizability, establishing strong benchmarks for vishing detection research.