Skip to content

Category

cybersecurity

1,065 papers

#machine learning Open access Sep 2026

ZK-Trace: Certified Collusion Tracing with Zero-Knowledge Credentials for Federated GNSS Interference Monitoring

Federated global navigation satellite system (GNSS) monitoring distributes a proprietary classifier to partly trusted stations, any of which may leak its copy. ZK-Trace combines public identity marks, recipient-specific Tardos fingerprints, and zero-knowledge credential verification. The registry supports offline traci...

Redwanul Karim, N. Raichur, Lucas Heublein et al. · 0 citations
#machine learning Preprint Sep 2026

When Topology Betrays Privacy: Lattice-Based Reconstruction Attacks on Secure Aggregation in Decentralized Federated Learning

The results show that colluding semi-honest nodes can recover the original local updates of honest nodes, enabling downstream reconstruction of private training data, and demonstrate that SA alone does not guarantee privacy in DFL when local aggregation induces asymmetric observations.

Wen-Rui Yu, Chang-Long Ji, Johannes Bjerva et al. · 0 citations
#machine learning Preprint Sep 2026

HoneyRoute: Honeypot-Model Routing for Adversarial LLM Serving

We introduce HoneyRoute, an inference-serving layer that detects whether an incoming request is malicious and, if so, routes it to a dedicated honeypot model, shielding production while the adversary's interaction is continuously harvested for intelligence. Existing defenses embed traps inside model memory or rebuild d...

Han Jin · 0 citations
#machine learning Preprint Open access Sep 2026

DRIFT: Removing Diffusion Watermarks by Deflecting the Generative Trajectory

Diffusion watermarking embeds verifiable signals into the generative process and commonly verifies them by recovering trajectory-dependent evidence, making the marks robust to conventional pixel-space distortions. Existing removal attacks either regenerate along deterministic trajectories, which often preserve the wate...

Rui Bao, Zheng Gao, Xiaoyu Li et al. · 0 citations
#machine learning Preprint Sep 2026

VEX-Bench: Benchmarking LLM Agents for Assessing Exploitability of Software Supply Chain Vulnerabilities

VEX-Bench is introduced, the first benchmark for evaluating LLM agents'ability to assess the exploitability of software supply chain vulnerabilities, and contains 75 real-world cases mined from GitHub and labeled by security experts, covering Python, Java, and Go.

Jia-Hao Shi, Edward Tsien, Yi-Feng Di et al. · 0 citations
#machine learning Review Sep 2026

Counter-Swarm Doctrine: Containing Coordinated Agent Intrusions

Agents can turn shared infrastructure into a channel for coordinated intrusion. The Hugging Face incident and a separate public-wiki investigation show why a security assessment may need evidence from several executions and the artifacts they leave behind. We argue that the operational unit of defence should be a revis...

Gregory N. Frank · 1 citation
#machine learning Preprint Open access Sep 2026

Topological Fraud Detection in Latent Transaction Spaces

Working entirely on topologically anonymized embeddings, we perform fraud detection using iterative rounds of unsupervised filtering followed by supervised sniping. The result is an ultra-low latency privacy--preserving triage that allows institutions to flag suspicious activity without compromising Personally Identifi...

Avraham Bourla · 0 citations
#machine learning Preprint Sep 2026

Robust Decentralized Personalized Federated Learning via Prediction-Constrained Neighborhood Collaboration

This paper proposes a robust decentralized personalized federated learning method R-DPFL, that enables clients to reduce the impact of Byzantine attacks via robust neighborhood direction estimation and history-based update trend prediction, rather than purely aggregating client models as in the existing work. In R-DPFL...

Xiao Ma, Hong Shen, Hui Tian et al. · 1 citation
#machine learning Preprint Sep 2026

The Oversight Gap: What LLM Safety Monitors Miss, and Why It Is Not Capability

Several properties safety monitors are asked to certify, among them cross-tenant noninterference, sandbagging and evaluation awareness, are 2-safety hyperproperties, witnessed only by two executions. The standard consequence is a binary impossibility: one trace cannot decide them. We replace the binary with a measureme...

Xin Xu · 0 citations
#machine learning Preprint Sep 2026

TrojanWorld: Backdooring World-Model Agents via Imagination Steering

To achieve effective, stealthy, and persistent control, TrojanWorld combines Decision-Reflective Induction to steer trigger-conditioned imagination toward attacker-specified actions using decision feedback, Clean Behavior Anchoring to preserve trigger-free predictive and behavioral fidelity, and Causal Propagation to s...

Wen-Kai Huang, Si-Yuan Liang, Gaolei Li et al. · 0 citations
#machine learning Review Sep 2026

MOLE: Detecting Insider Threats in AI Agents

Model misalignment, prompt injection, or operator misuse could lead AI agents operating frontier-lab accounts to exfiltrate model weights, poison training data, or weaken release gates. Existing benchmarks do not test whether defenders can detect this activity among routine work under a limited review budget. We introd...

Aashiq Muhamed, Virginia Smith · 0 citations

From tech blogs

See all →
Google DeepMind Blog Jul 17, 2026

Introducing Gemini 3.5 Flash Cyber

Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.