Findings indicate that models optimized for local n-gram patterns provide a highly cost-effective deployment solution while outperforming representative pre-trained language models in specialized adversarial digital forensics tasks.
Abstract
Drug-related criminal activities on social media increasingly employ rapidly evolving coded languages, including fruit substitutions, numeric homophones, and dialectal metaphors, to evade detection. This adversarial obfuscation causes large-scale deep learning models to suffer from severe computational overhead during edge deployment, while reducing their robustness against evolving coded expressions. To mitigate these challenges, we construct a dedicated dataset containing 10,000 drug-related coded text samples and propose an optimized, lightweight TextCNN-based framework. The framework normalizes lexical variants of codewords using a domain-specific dictionary and adaptive normalization functions, and extracts local semantic patterns from the word embedding layer through multi-scale convolutional kernels (h∈{3,4,5}) to capture crucial short-text semantics in parallel. Experimental results demonstrate that the proposed framework achieves an F1 score of 99.3% with only 0.22 M parameters, significantly outperforming baseline models. These findings indicate that models optimized for local n-gram patterns provide a highly cost-effective deployment solution while outperforming representative pre-trained language models in specialized adversarial digital forensics tasks.
This study demonstrates how integrating contextual cues and custom lexical signals can significantly improve the detection of AI-generated phishing content and develop sophisticated and resilient defenses against emerging AI-enabled threats.
Raghad Ghawa, A. Alhogail· Electronics· 0 citations
A feature-driven framework for phishing Uniform Resource Locator (URL) detection is introduced, emphasizing the design and evaluation of enhanced feature representations and highlighting that performance gains are primarily driven by feature design rather than model complexity.
Deniz Kaya, Murat Osmanoğlu· PeerJ Computer Science· 0 citations
A hybrid detection framework which combines semantically deep embeddings from the RoBERTa transformer with a set of carefully designed language statistics and linguistic statistics and shows excellent resistance to the surface-level adversarial paraphrasing strategy.
Anita Rani, Suman· International Journal of Sci...· 0 citations
The rising prevalence of substance abuse and overdose incidents underscores the need for real-time public health surveillance. Social media offers valuable signals for monitoring these events; however, noisy language, slang usage, and class imbalance present significant challenges for automated analysis. To address these issues, the authors propose ATTEND, a multi-task neural network for substance classification and detection of 18 overdose symptoms, with symptom normalization to standardized MedDRA concepts. ATTEND was trained on a large multi-source corpus combining ADE Corpus V2 and the UCI Drug Review Dataset, comprising over 100,000 samples designed to emulate realistic social media communication. Experimental results show that ATTEND achieved 93.23% accuracy and 93.41% weighted-F1 for substance classification, 94.10% micro-F1 for overdose symptom detection, and 90.42% accuracy for symptom normalization, outperforming baseline multi-task models across all tasks. The framework is scalable, privacy-preserving, and suitable for real-time monitoring of drug abuse signals.
Sudhakar Kumar, Sunil K. Singh, Satvik Pathak et al.· International Journal of Int...· 0 citations
Cyberbullying via social media is a constant digital safety issue because the content can be widely shared and openly visible and can have a negative impact on users before it is removed by manual moderation. Current detection models are mostly based on shallow lexical features or transformer-only classifiers, resulting in low-level accuracy and explainability. This study introduces a Hybrid BERT–XGBoost model to detect cyberbullying in short social media texts, which combines the strengths of both models. The contextual sentence embeddings are extracted using BERT and the auxiliary linguistic and behavioral features are extracted in parallel, such as sentiment polarity, profanity score, punctuation intensity, capitalization ratio, hashtag usage, mention count, emoji frequency, and post length. XGBoost is used for the classification of the fused representation. The model is tested on stratified training, validation, and testing splits, compared to a baseline model, ablated, tested with macro-F1, weighted-F1, ROC-AUC, early detection recall, and grouped explainability. The proposed framework achieved 96.18% accuracy, 96.05% macro-F1, 96.16% weighted-F1, 95.88% early detection recall, and 98.42% macro-AUC. It performs better than the BERT + Dense baseline, which obtained 94.31% accuracy and 94.08% macro-F1 score, demonstrating the advantage of fusion of contextual and auxiliary features. The framework provides an interpretable, practical and category-aware solution for early detection of cyberbullying, but further research is needed to validate the framework in multiple languages, modalities and in conversations.
Rima Shishakly, Abdelhadi Mohammad Bayoud, Mansour Obeidat et al.· Journal of Cyber Security an...· 0 citations
PhishingGAT, a detector that fuses word-level semantic features with structural ones and is hardened against adversarial perturbation, is presented, a detector that fuses word-level semantic features with structural ones and is hardened against adversarial perturbation.
R. Kodali, Siva Rama Krishna T Dr· International Journal of Inn...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.