Skip to content

Category

cybersecurity

1,065 papers

#machine learning Preprint Open access Oct 2026

On Reliability of Membership Inference Vulnerability Evaluation

Membership inference attacks (MIAs) are popular methods for empirically assessing the leakage of sensitive information in the training data through models or statistics learned from the data. The MI vulnerability is often evaluated through a binary classifier that tries to predict whether a particular sample was in the...

Joonas J\"alk\"o, Gauri Pradhan, Ossi R\"ais\"a et al. · 0 citations
#machine learning Preprint Oct 2026

Learnt Attacks on Quantum Key Distribution under Channel Noise and Device Drift

Quantum key distribution (QKD) links are provisioned from security analyses of stationary channels, whereas the devices that determine the channel drift between recalibrations. Whether an eavesdropper who cannot alter the channel's own noise gains by following that drift has not been quantified. Adaptive eavesdropping...

Marcel Mordarski, Benjamin I. Gräs, A. Shehata et al. · 0 citations
#machine learning Preprint Oct 2026

The Innocent Courier: Covert Exfiltration Through Legitimate LLM Web Fetching

With the increasing capabilities of Large-Language-Models (LLMs) and LLM-based agents, users are increasingly using them to solve everyday problems, such as answering e-mails or providing programming support. Existing work has extensively investigated security and privacy risks, such as prompt injections and the disclo...

Alessandro Pegoraro, Daryan Merx, P. Rieger et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Crossing the Cyber Divide: Sim-to-Sim and Sim-to-Real Transfer for RL Agents

Cyber attack agents are typically trained and evaluated within a single simulator, making it unclear whether learned policies transfer beyond the environments in which they were developed. This limitation hinders both deployment and fair comparison, as cyber simulators differ substantially in their state representation...

Sabrina Saika, Yinuo Du, Aritran Piplai · 0 citations
#machine learning Preprint Open access Oct 2026

Progressive-Resolution Secure Aggregation for Federated Learning

Secure aggregation lets a server recover an aggregate of client updates without observing any individual update, but conventional protocols fix the aggregate precision when clients upload. We introduce and formulate a new progressive-resolution secure-aggregation functionality in which clients upload once and successiv...

Seyed Mohammad Azimi-Abarghouyi · 0 citations
#machine learning Preprint Open access Oct 2026

On the Relationship between Model Quantization and Model Inversion Attacks

Model quantization reduces the numerical precision of neural network weights and activations to lower storage and computational costs. Model inversion attacks recover or reconstruct sensitive training data or inference inputs from model outputs or intermediate features, so quantization may also alter their effectivenes...

Rongke Liu, Youwen Zhu · 0 citations
#machine learning Preprint Open access Oct 2026

Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisation

Backdoor attacks can be implanted in Large Language Models (LLMs) during training, causing unwanted behaviour when a trigger appears in the input. Existing backdoor defences for LLMs attempt to remove the backdoor but inadvertently shift the model's output distribution to benign prompts, which can result in degraded mo...

Minoo Kim, Vasileios Lampos, George Drayson · 0 citations
#machine learning Preprint Open access Oct 2026

Refusal Localizes, the Damage Relocates: Safety Layers Under Few-Sample Fine-Tuning

Fine-tuning adapts aligned large language models (LLMs) to downstream tasks, but a few dozen harmful examples can remove their refusal of harmful requests. Prior work localizes safety-related behavior to specific layers, directions, and tokens, suggesting targets for protection. We test whether successful localization...

Jungseob Lee, Dongyub Jude Lee, Sugyeong Eo et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Evasion Attacks: How Adversarial Noise Bypasses ML Classifiers

This paper presents a reproducible, educational study of evasion attacks in image classification and text classification. A compact convolutional network trained on MNIST reached 98.63% clean test accuracy and was evaluated under two white-box attacks. Under FGSM, accuracy fell to 60.20% at $\epsilon$ = 0.15 and 1.72%...

Parker Hummel (Minot State University), Ryne Skabo (Minot State University), Muhammad Abusaqer (Minot State University) · 0 citations
#machine learning Preprint Oct 2026

A Safe Prototype Is Not a Safety Direction: Reference Dependence and Prompt Confounds in Response-Safety Embeddings

Can response safety be scored by cosine similarity to the mean embedding of known-safe responses? A recent sleeper-agent detector proposes exactly this score, yet the raw positive-centroid rule is not identified: positive observations locate the safe class relative to an encoder origin, but do not determine which direc...

Sahil Kadadekar · 0 citations
#machine learning Preprint Open access Oct 2026

SAGE: Similarity-Based Cleaning of Poisoned Training Data from Verified Examples

As machine learning increasingly relies on public, untrusted data sources, data poisoning attacks, which inject malicious examples into training data to induce misclassification of a chosen target, pose a growing threat. Existing defenses either assume zero ground-truth information about which examples are poisoned, or...

Chaeeun Han, Soodeh Atefi, Yevgeniy Vorobeychik et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

TensorCommitments: A Lightweight Verifiable Inference for Language Models

Most large language models (LLMs) run on external clouds: users send a prompt, pay for inference, and must trust that the remote GPU executes the LLM without any adversarial tampering. We critically ask how to achieve verifiable LLM inference, where a prover (the service) must convince a verifier (the client) that an i...

Oguzhan Baser, Elahe Sadeghi, Eric Wang et al. · 0 citations

From tech blogs

See all →
Google DeepMind Blog Jul 17, 2026

Introducing Gemini 3.5 Flash Cyber

Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.