Experiments on certain machine learning models and two transformers were conducted to detect unstructured, textual, context-dependent sensitive data and show that DistilRoberta demonstrated higher accuracy and recall, and was faster and lighter than ALBERT.
Abstract
The massive amount of publicly available data has necessitated an increase in public and organizational awareness of the potential risks of leaking private data, whether intentionally or unintentionally. The damage caused by leaking these data depends on their degree of sensitivity. Disclosing a person’s or an organization’s private data via different social media platforms might threaten people’s lives or the organization’s reputation or finances. Handling big data, especially unstructured data, is challenging. Consequentially, many solutions have been proposed to detect sensitive data in structured containers. However, detecting sensitive data in unstructured containers is still challenging, especially with context-dependent and high-performance measurement results. In this study, experiments on certain machine learning models and two transformers—DistilRoberta and ALBERT—were conducted to detect unstructured, textual, context-dependent sensitive data. The results show that DistilRoberta demonstrated higher accuracy and recall, and was faster and lighter than ALBERT.
A systematic evaluation of several proprietary and open-source Large Language Models for sensitive entity extraction from documents and a form-filling pipeline that uses vision-capable LLMs to label form fields, generate realistic synthetic personas, and fill real blank forms, enabling reproducible evaluation across di...
Errita Xu, Stefan Larson, Kevin Leach· Proceedings of the 2026 ACM...· 0 citations
This paper introduces the development path of sensitive data detection technology, including rule-based methods, older machine-learning techniques, and currently popular pre-trained language models, and introduces the research content of LLM-driven zero-shot named entity recognition, open information extraction and pri...
Jie-Qun Wei, Yuejin Zhang· Scientific Journal of Intell...· 0 citations
The findings in this study highlight the potential and current limitations of LLMs for phishing detection in dynamic instant messaging environments and emphasize the superior performance of platform-tailored models.
Md Erfan, Paula Branco, Guy-Vincent Jourdan· Digital Threats: Research an...· 0 citations
Phishing remains one of the most persistent cyber threats, particularly in email environments where deceptive messages can be distributed at scale. This paper compares five classifiers: Multinomial Naive Bayes, Random Forest, Bidirectional Long Short-Term Memory (BiLSTM), DistilBERT, and BERT-base. A multi-source corpu...
Andre Sebastian Samaniego Buñay, Ariel Misael Orellana Albarracin, Joel Marcelo Chuquimarca Pomagualli· Enfoque UTE· 0 citations
Abstract—. Privacy-preserving data publishing has become vital in the age of data-driven decision-making, especially with the increasing availability of unstructured datasets. While the Score, Arrange, and Cluster (SAC) algorithm effectively anonymizes structured data, it does not address the challenges posed by unstru...
K. G, Kiran B. M., G. Prasad· International Scientific Jou...· 0 citations
Despite the advancements made by researchers, spam emails remain one of the biggest challenges in the field of cybersecurity. Spam emails can serve as phishing emails or carry viruses that compromise the security of an organization's system. Current detection techniques depend on supervised learning or rely on cloud-ba...
Vusal Shahbazov· 2026 7th International Confe...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.