Skip to content

Category

cybersecurity

1,032 papers

#machine learning Preprint Open access Oct 2026

Closed-Form Noise Calibration Against Membership Inference for Random-Allocation DP-SGD

DP-SGD protects training data by adding Gaussian noise to clipped gradients. The amount of noise is usually chosen by running a numerical privacy accountant inside a search. We study DP-SGD with random allocation, where each epoch uses every record once, at a randomly chosen step. For this setting we give a one-line fo...

Murat Bilgehan Ertan, Marten van Dijk · 0 citations
#machine learning Preprint Open access Oct 2026

Removing Information Content Does Not Certify Tamper Resistance in Open-Weight Models

Does removing harmful information make open-weight models resistant to fine-tuning attacks? We show that mutual information at release alone cannot universally certify slow recovery. Function-preserving reparameterizations leave information unchanged while altering gradient-descent geometry, so an invariant certificate...

Domenic Rosati, Alessa Carbo, Ali Dadsetan et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Curvature-Guided Module Localization for Low-Rank Detoxification of Backdoored Large Language Models

Backdoor attacks pose a serious threat to large language models (LLMs) by causing otherwise benign systems to produce attacker-specified malicious behavior when a hidden trigger is present. In this work, we study post hoc detoxification of backdoored LLMs in a practical setting where the defender has access to the pois...

Arash Raftari, Mehrdad Mahdavi, Nathan Blackthorn et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

ArapaiSecure: An Autonomous AI Security Agent for Banking: Multi-Vector Fraud and AML Detection Across Retail and Corporate Accounts

Banks face two threat families with fundamentally different detection requirements: signature-based fraud (card- not-present attacks, account takeover, ATM cloning) and behavioural financial crime (structuring, layering, mule networks, business email compromise). Static rule engines catch high-velocity events but remai...

Joseph Walusimbi, Joshua Benjamin Ssentongo · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Hidden in Plain Sight: Benchmarking Agent Safety Against Decomposition Attacks with DECOMPBENCH

LLM-based Agents are becoming increasingly capable and widely deployed, creating growing incentives for adversarial misuse in the real-world. A key emerging threat is Decomposition Attacks \cite{glukhov2024breach, jones2024adversaries} in which a harmful task is broken into simpler, benign subtasks that evade safety me...

Vikhyath Kothamasu, Virginia Smith, Chhavi Yadav · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Targeting World Models to Compromise Robot Learning Pipelines

World models have recently seen a rapid growth in both their popularity and capability as more data efficient tools for generating robot training data or simulating real world environments, with many works proposing their integration into the robot learning pipeline. While highly practical, in this work we demonstrate...

Ethan Rathbun, Ahmed Agha, Saaduddin Mahmud et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

A Few Steps Further: Why Defenses Against Malicious Finetuning Erode Under Continued Training

Model providers increasingly release the weights of large language models. Although these models are safety-aligned before release, their safeguards can often be removed by fine-tuning on harmful data. A growing class of defenses aims to make alignment robust to such malicious fine-tuning, but these defenses are typica...

Itay Zloczower, Eyal Lenga, Gilad Gressel et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

A Survey of Secure Retrieval-Augmented Generation

Retrieval-augmented generation (RAG) improves large language models (LLMs) with external knowledge, but this access path creates security risks distinct from inherent prompt-only or parametric-model flaws. We frame secure RAG as securing external knowledge access. We conducted a systematic search and curated 135 works...

Yuming Xu, Mingtao Zhang, Zhuohan Ge et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

SLDR: Defending Against Malicious Fine-tuning via Selective Layers Recovery and Dynamic Routing

Fine-tuning-as-a-service enables users to adapt aligned large language models (LLMs) to specialized tasks, but malicious fine-tuning can erode refusal behavior while preserving task performance on legitimate inputs. We revisit recent layer-wise safety diagnostics and find that safety sensitivity is signed: scaling diff...

Hui Zhang, Yachao Yuan, Jiayun Wang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Constrained-Action AI Remediation for SIEM/XDR via a NeMo-Guardrails Proxy

Security Operations Centers (SOCs) for information technology and operational technology share one incident-response problem: a flood of correlated alerts and too few analysts. Large Language Models (LLMs) are increasingly proposed as reasoning engines that triage alerts and, in autonomous deployments, issue commands t...

Georgios Koutidis, Nikolaos Kekatos, Tom Nianios et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Defensive Sufficiency in a Stackelberg Model of AI Security

Feedback from automated testing, human red teaming, and incident response can strengthen an AI system's defenses when discovered failures lead to effective repairs. We study when this feedback process provides sufficient protection and when investing in it is economically worthwhile. We begin by showing that an attack...

Subhabrata Majumdar, Rajlakshmi Chavan · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Correct Answers, Unsupported Findings: Evidence Binding in Forensic Reconstruction of LLM Agent Logs

Forensic reconstruction of LLM-agent actions requires not only recovering the correct value, but establishing which preserved record supports that finding. Tool logs, generated explanations, and local citation identifiers capture different parts of this evidence, yet a citation identifier does not establish a source un...

Taehyeon Yun, Dongho Kim, Geonwoo Kim et al. · 0 citations

From tech blogs

See all →
Google DeepMind Blog Jul 17, 2026

Introducing Gemini 3.5 Flash Cyber

Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.