Skip to content

Category

cybersecurity

1,065 papers

#natural language process... Preprint Open access Oct 2026

The Same Zero: Why Identical ASR Can Imply Different Guarantees in LLM-Agent Security

LLM-agent security has produced a dense landscape of defenses - prompt hardening, content filters, permission gates, sandboxes - yet no framework tells a deployer what a defense actually guarantees, or where that guarantee comes from. We apply Verification Autonomy Levels (VAL) - L0: LLM self-declaration; L1: determini...

YaJie Yin · 0 citations
#natural language process... Preprint Open access Oct 2026

Large Language Models and Augmented Democracy

Artificial intelligence enables computational agents to represent political preferences and take part in collective decision-making. In this thesis, I investigate the opportunities and challenges of digital twins (DTs) based on Large Language Models (LLMs) as intermediaries in augmented democracy, focusing on individua...

Jairo Gudi\~no-Rosero · 0 citations
#natural language process... Preprint Open access Oct 2026

Red-TTT: Test-Time Training for Automated Jailbreaking Large Language Models

Large language models remain vulnerable to jailbreaks, and automated red teaming is the standard way to find jailbreaks in large language models at scale. Current methods either draw more samples at test time through search, rewriting, and tree expansion, or train a stronger attacker offline with reinforcement learning...

Tongyan Hu, Hao Li, Xiaogeng Liu et al. · 0 citations
#machine learning Preprint Open access Oct 2026

EXHIB: A Benchmark for Realistic and Diverse Evaluation of Function Similarity in the Wild

Binary Function Similarity Detection (BFSD) is a core problem in software security, supporting tasks such as vulnerability analysis, malware classification, and patch provenance. In the past few decades, numerous models and tools have been developed for this application; however, due to the lack of a comprehensive univ...

Yiming Fan (The Ohio State University), Jun Yeon Won (The Ohio State University), Ding Zhu (The Ohio State University) et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Logit-Gap Steering: A Forward-Pass Diagnostic for Alignment Robustness

RLHF-style alignment trains language models to refuse unsafe requests, but how much operational margin does this refusal rest on? We introduce the refusal-affirmation logit gap: the difference between the top refusal-token logit and the top affirmative-token logit at the first decoding step. This single scalar quantifi...

Tung-Ling Li, Hongliang Liu · 0 citations
#machine learning Preprint Open access Oct 2026

Backdoor Attacks on Discrete Graph Diffusion Models

Diffusion models have demonstrated remarkable generative capabilities in continuous data domains such as images and videos. Recently, discrete graph diffusion models (DGDMs) have extended this success to graph generation, achieving state-of-the-art performance. However, deploying DGDMs in safety-critical applications,...

Jiawen Wang, Samin Bin Karim, Yuan Hong et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Your Privacy My Cloak: Backdoor Attacks on Differentially Private Federated Learning

Prior research suggests that differential privacy (DP) can enhance the robustness of federated learning (FL) against backdoor attacks. In this paper, we challenge this assumption. Through an empirical analysis of two baseline attack strategies, we uncover a fundamental tension in DP-FL: while bypassing DP allows state-...

Xiaolin Li, Ning Wang, Ninghui Li et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Sample-wise Targeted Adversarial Attacks on Test-time Adaptation

Test-time adaptation (TTA) mitigates distribution shifts by adapting models to unlabeled test inputs, but also exposes them to adversarial manipulation. Existing class-wise targeted attacks remain suboptimal for stealthy exploitation in this setting: since TTA operates on batches, forcing a subset of samples toward a t...

Phuc Duc Nguyen, Quang Duc Nguyen · 0 citations
#machine learning Preprint Open access Oct 2026

NASimJax: A GPU-Accelerated Policy Learning Framework for Penetration Testing

Penetration testing - the practice of simulating cyberattacks to identify vulnerabilities - is a complex sequential decision-making task that is inherently partially observable and features large action spaces. Existing RL simulators for this domain are CPU-bound and fixed to narrow scenarios, making it infeasible to t...

Raphael Simon, Jos\'e Carrasquel, Elli Makdis Antoun et al. · 0 citations
#machine learning Preprint Open access Oct 2026

How Vulnerable Is My Learned Policy? Universal Adversarial Perturbation Attacks On Modern Behavior Cloning Policies

Imitation learning, also known as learning from demonstrations, is a popular approach to train AI models; however, the vulnerability of these models to adversarial attacks remains underexplored. We present the first systematic study of adversarial attacks, across a range of both classic and recently proposed imitation...

Akansha Kalra, Basavasagar Patil, Guanhong Tao et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Resisting Adversarial Attacks in Deep Neural Networks using Diverse Decision Boundaries

The security of deep learning (DL) systems is an extremely important field of study as they are being deployed in several applications due to their ever-improving performance to solve challenging tasks. Despite overwhelming promises, the deep learning systems are vulnerable to crafted adversarial examples, which may be...

Manaar Alam, Shubhajit Datta, Debdeep Mukhopadhyay et al. · 0 citations
#machine learning Preprint Oct 2026

Correct Verdicts, Flawed Reasoning: Structured Auditing of LLM-based Vulnerability Reasoning

Large Language Models (LLMs) are increasingly deployed for automated software vulnerability analysis. Binary classification alone is insufficient; practitioners need explanations to triage bugs and engineer patches. Standard practice relies on Chain-of-Thought (CoT) prompting, but free-form reasoning allows models to o...

Boyue Caroline Hu, K. Ahir, Ronghao Ni et al. · 0 citations

From tech blogs

See all →
Google DeepMind Blog Jul 17, 2026

Introducing Gemini 3.5 Flash Cyber

Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.