Skip to content

Category

cybersecurity

1,065 papers

#machine learning Preprint Sep 2026

Decoy Direction Optimization: A Post-Hoc Defense Against LLM Abliteration

Decoy Direction Optimization (DDO), a fast, post-hoc weight-editing defense that requires no base-model finetuning, is introduced, based on a simple mechanistic insight: ablation attacks rely on contrastive estimators to find the refusal direction.

Aashiq Muhamed, Mona T. Diab, Virginia Smith · 2 citations · ⚡1
#machine learning Preprint Sep 2026

SWB-DM: A Calibrated Sliced-Wasserstein-Barycenter Aggregator with Delayed-Momentum Caching for Byzantine-Robust Federated Learning under Partial Participation

SWB-DM's cache carries a real warm-up cost, but extending all baselines to the same round budget shows its CIFAR-10 gains are disproportionately large, and on CIFAR-100, FLTrust benefits more -- for reasons entirely unrelated to caching.

S. Saranraj, S. SaranyaM, S. AlexDavid et al. · 0 citations
#cybersecurity Preprint Open access Sep 2026

Keys on Doormats: Exposed API Credentials on the Web

API (Application Programming Interface) keys allow applications to authenticate themselves to third-party services. Inadvertent public exposure of these credentials can pose significant consequences, as adversaries can use them to gain privileged access to other services. In this paper, we measure API credential exposu...

Nurullah Demir (Stanford University), Yash Vekaria (University of California, Davis) et al. · 0 citations
#cybersecurity Preprint May 2025

Towards the ideals of Self-Recovery and Metadata Privacy in Social Vault Recovery with Apollo

Apollo is a social recovery mechanism that aims to avoid any memorability assumptions while strongly protecting recovery metadata privacy, and uses a novel multi-layered secret sharing scheme to mitigate the computational overhead of recovery in this setting, which would otherwise be exponential in the recovery thresho...

Shailesh Mishra, Simone Colombo, Pasindu Tennage et al. · 0 citations

PARSE: Provenance-Aware Retrieval Sanitization for Professional Domain LLM Agents

This work introduces PARSE (Provenance-Aware Retrieval Sanitization), a domain-aware, fact-preserving sanitization pipeline that classifies each sentence by injection likelihood, extracts structured facts before rewriting, and verifies fact preservation via a consistency-checking loop.

Aaditya Pai · 0 citations
#natural language process... Preprint Sep 2026

CiteShade: Citation Laundering in Multi-Source Retrieval-Augmented Generation and Its Counterfactual Defense

This work proposes CiteShade, the first citation laundering attack to RAG, in which an attacker controlling a single source induces a model to produce an attacker-chosen wrong answer and to attribute it to a trusted source that does not support it, while the evidence for the correct answer remains in context.

Fu-Zheng Guo · 0 citations

Differential Privacy of Gaussian Process Posterior Sampling

This work derives Renyi-DP guarantees separating privacy leakage through the posterior mean from a distinct channel induced by the data-dependent posterior covariance, and identifies effective ridge regularisation and covariance scale as the principal privacy-controlling quantities.

Tomasz Maciazek · 0 citations
#machine learning Preprint Open access Sep 2026

Noise-Aware and Dynamically Adaptive Federated Defense Framework for SAR Image Target Recognition

As a critical application of computational intelligence in remote sensing, deep learning-based synthetic aperture radar (SAR) image target recognition facilitates intelligent perception but typically relies on centralized training, where multi-source SAR data are uploaded to a single server, raising privacy and securit...

Yuchao Hou (Shanxi Normal University, Taiyuan, China) et al. · 0 citations
#machine learning Preprint Open access Sep 2026

Fusing Spectral Signatures and Activation Clustering for Backdoor Detection in Healthcare Imaging Models: Method, Implementation, and Evaluation

Machine learning models are increasingly deployed in healthcare imaging pipelines for diagnostic support, and training-time attacks against them are a named sector-level concern: healthcare-sector guidance identifies model poisoning and adversarial attacks as threats requiring dedicated defenses, while federal policy d...

Suresh Tamang · 0 citations
#machine learning Preprint Sep 2026

EI-DDLGN: Efficient Encrypted Inference with Deep Differentiable Logic Gate Networks under TFHE

EI-DDLGN is presented, the first in-depth study of TFHE-based DDLGN inference, and how encrypted execution cost depends on model size, learned Boolean-function distribution, and propagated wire status is characterized, and Model-Fixed-Wire PBS Bypass (MFW-PBS Bypass), a semantics-preserving execution strategy that elim...

Mahmoud Y. M. Yassin, Mahmoud Abdelhafeez Sayed, Mostafa Taha · 0 citations
#artificial intelligence Preprint Open access Sep 2026

DualView: Preventing Indirect Prompt Injection in Personal AI Agents

Personal AI agents that run on the user's local machine automate daily tasks including web search, email, and file management. Their access to computer resources, including the network, file system, and shell, exposes them to indirect prompt injection (IPI) attacks. Prior Dual LLM defenses block IPI by replacing untrus...

Juhee Kim, Woohyuk Choi, Taehyun Kang et al. · 0 citations

From tech blogs

See all →
Google DeepMind Blog Jul 17, 2026

Introducing Gemini 3.5 Flash Cyber

Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.