Jul 2026· 2026 3rd World Conference on Computer and Information Security (WCCIS)· pp. 231-236· 0 citations· 14 references
Abstract
Malicious URL detection serves as a critical component in safeguarding cyberspace security. Although deep learning models have achieved remarkable performance in identifying complex malicious patterns, their large parameter size and high computational overhead have become bottlenecks restricting their real-time deployment in resource-constrained environments such as gateways and mobile terminals. To address this challenge, this paper proposes Purl-Distill, a lightweight malicious URL detection framework based on multi-granularity knowledge distillation. Different from traditional compression methods that only focus on output alignment, Purl-Distill introduces a multi-granularity alignment mechanism. First, by constructing a multi-layer feature mapping space, it deeply transfers the hierarchical representations of the teacher model from low-level character morphology to high-level semantic logic. Second, a temperature-scaled logits distribution alignment strategy and a joint loss function are designed to accurately capture the decision boundary features of the teacher model. Experimental results demonstrate that on multiple public benchmark datasets, Purl-Distill achieves an order-of-magnitude reduction in model parameters and storage space while maintaining detection accuracy comparable to state-of-the-art (SOTA) teacher models. This work provides an efficient engineering solution for high-performance network security protection in resource-constrained scenarios.
A hierarchical binary classification framework that decomposes the multiclass task into a sequence of structured binary decisions and provides competitive and feature-efficient classification performance is proposed.
Mehmet Batuhan Ozdas, Murat Osmanoglu· Journal of Computer Virology...· 0 citations
Malicious URLs are evolving with increasingly sophisticated obfuscation techniques, posing a persistent threat to global cybersecurity. Traditional deep learning methods still struggle to identify the intricate topological characteristics and hidden structural dependencies inherent in URL sequences. This paper proposes a structural-semantic integration framework leveraging Graph Attention Networks and Knowledge Distillation for robust multi-layer malicious URL detection. We conceptualize raw URLs as character-level directed graphs to extract intricate structural patterns, fused with high-level lexical context distilled from a pretrained DistilBERT model. A hybrid loss integrating Kullback-Leibler divergence and Negative Log-Likelihood balances computational efficiency and predictive performance. Experimental results on large-scale, imbalanced datasets show our framework achieves 98.86% accuracy and 98.16% F1-score.
Hieu Nguyen Trung, Cường Phạm Việt, Ngoc Tran Nguyen· Journal of Military Science...· 0 citations
A hybrid phishing detection framework that integrates three complementary techniques: DistilBERT for semantic analysis of URL text, Graph Neural Networks for modelling structural relationships among URL components, and LightGBM for efficient metadata-based feature classification is proposed.
I. Shalini, G. Sujini· International Journal for Re...· 0 citations
A high-efficiency detection framework utilizing DistilBERT, a distilled knowledge representation of the BERT transformer is proposed, substantiate the viability of Knowledge Distillation as a mechanism to deploy state-of-the-art semantic security filters on edge infrastructure.
Mrinal Mrinal, Neeraj Kumar· International Journal of Cre...· 0 citations
The proliferation of algorithmically generated malicious URLs presents a critical challenge for modern cybersecurity, requiring a shift from syntactic pattern-matching toward semantic understanding grounded in pre-trained transformer foundations. Building on our prior work on billion-scale semantic search and density-based campaign clustering, this paper presents a unified, deployable framework for real-time campaign discovery. The framework converts raw URL streams into dense Sentence-BERT embeddings and couples approximate nearest neighbor search with online density-based clustering, discovering emerging campaigns without prior knowledge of their number or shape. Our central finding is that the choice of semantic representation is decisive: a domain-focused embedding strategy yields near-perfect campaign separation, substantially outperforming full-URL representations. On live, in-the-wild threat feeds, the domain-focused representation recovers all 944 discovered campaigns at an Adjusted Rand Index of 0.990 and a mean campaign recall of 1.000, at 0.10 ms per URL. Under identical clustering, the full-URL representation reaches an Adjusted Rand Index of only 0.510 and recovers fewer than half of the campaigns. We add a SHAP-based explainability layer that attributes discovery decisions to interpretable structural patterns, and we expose the whole pipeline as an operational system with a REST interface for single- and batch-URL analysis. The encoder we employ is a discriminative representation model rather than a generative one; what we contribute is the representation-and-modeling discipline this setting demands, and the resulting guidance transfers to threat intelligence in an era of generatively produced attack campaigns.
G. Feretzakis, Dimitrios Karapiperis, S. Mitropoulos· Information· 0 citations
GAFPNet (Generalization-Aware and False-Positive Controlled Framework Network), a five-module stacked-ensemble framework, is introduced as a generalization-aware and deployment-oriented phishing URL classifier for real-time filtering applications.
Mohammed Elias Basha S., M. N.· Journal of Trends in Compute...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.