Skip to content
Open access

Detecting Context-Dependent Sensitive Data in Unstructured Text

Jul 2026 · Information · Vol 17, pp. 663 · 0 citations · 53 references

TL;DR

Experiments on certain machine learning models and two transformers were conducted to detect unstructured, textual, context-dependent sensitive data and show that DistilRoberta demonstrated higher accuracy and recall, and was faster and lighter than ALBERT.

Abstract

The massive amount of publicly available data has necessitated an increase in public and organizational awareness of the potential risks of leaking private data, whether intentionally or unintentionally. The damage caused by leaking these data depends on their degree of sensitivity. Disclosing a person’s or an organization’s private data via different social media platforms might threaten people’s lives or the organization’s reputation or finances. Handling big data, especially unstructured data, is challenging. Consequentially, many solutions have been proposed to detect sensitive data in structured containers. However, detecting sensitive data in unstructured containers is still challenging, especially with context-dependent and high-performance measurement results. In this study, experiments on certain machine learning models and two transformers—DistilRoberta and ALBERT—were conducted to detect unstructured, textual, context-dependent sensitive data. The results show that DistilRoberta demonstrated higher accuracy and recall, and was faster and lighter than ALBERT.

Read PDF

Similar papers

Book Open access Aug 2026

Sensitive Data Detection in Documents with LLMs

A systematic evaluation of several proprietary and open-source Large Language Models for sensitive entity extraction from documents and a form-filling pipeline that uses vision-capable LLMs to label form fields, generate realistic synthetic personas, and fill real blank forms, enabling reproducible evaluation across di...

Errita Xu, Stefan Larson, Kevin Leach · 0 citations
Review Open access Aug 2026

A Survey of Zero-Shot Sensitive Information Detection Techniques based on Large Language Models

This paper introduces the development path of sensitive data detection technology, including rule-based methods, older machine-learning techniques, and currently popular pre-trained language models, and introduces the research content of LLM-driven zero-shot named entity recognition, open information extraction and pri...

Jie-Qun Wei, Yuejin Zhang · 0 citations
Open access Jul 2026

Can LLMs Keep Up? Evaluating Phishing Detection on Telegram

The findings in this study highlight the potential and current limitations of LLMs for phishing detection in dynamic instant messaging environments and emphasize the superior performance of platform-tailored models.

Md Erfan, Paula Branco, Guy-Vincent Jourdan · 0 citations
Open access Sep 2026

Comparative Analysis of Transformer-Based and ClassicalMachine Learning Models for Phishing Email Detection:A Multi-Source Dataset Evaluation with Explainability

Phishing remains one of the most persistent cyber threats, particularly in email environments where deceptive messages can be distributed at scale. This paper compares five classifiers: Multinomial Naive Bayes, Random Forest, Bidirectional Long Short-Term Memory (BiLSTM), DistilBERT, and BERT-base. A multi-source corpu...

Andre Sebastian Samaniego Buñay, Ariel Misael Orellana Albarracin, Joel Marcelo Chuquimarca Pomagualli · 0 citations
Jul 2026

Enhanced SAC For Text Privacy

Abstract—. Privacy-preserving data publishing has become vital in the age of data-driven decision-making, especially with the increasing availability of unstructured datasets. While the Score, Arrange, and Cluster (SAC) algorithm effectively anonymizes structured data, it does not address the challenges posed by unstru...

K. G, Kiran B. M., G. Prasad · 0 citations
Conference Jul 2026

Performance and Explainability of Open-Weight Large Language Models for Spam Email Detection

Despite the advancements made by researchers, spam emails remain one of the biggest challenges in the field of cybersecurity. Spam emails can serve as phishing emails or carry viruses that compromise the security of an organization's system. Current detection techniques depend on supervised learning or rely on cloud-ba...

Vusal Shahbazov · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.