Skip to content

Category

cybersecurity

1,065 papers

GREAT: Generalizable Backdoor Attacks in RLHF via Emotion-Aware Trigger Synthesis

This work develops GREAT, a novel framework for crafting natural distributional backdoors in RLHF, which targets harmful response generation for a vulnerable user subpopulation featured by semantically violent requests paired with emotionally angry triggers.

Subrat Kishore Dutta, Yuelin Xu, P. Pant et al. · 0 citations
#machine learning Preprint Aug 2026

REPLICANT: Learning Policies for Evading and Hardening Malware Detectors

This work presents Replicant, a deep reinforcement learning framework that learns the realistic task of evasion under a strict label-only black-box threat model and demonstrates that learning the task of evasion not only results in stronger attack performance but provides a better signal for hardening malware detectors...

Shae McFadden, Ilias Tsingenopoulos, Mario D'Onghia et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents

LoopHarness is presented, which restores a persistent, non-decaying safety state at the loop level at the loop level, and gives a complete evaluation protocol on native Agent-SafetyBench tasks with paired clean and attacked episodes, an outer-state attack suite whose decisive evidence exists only across iterations, per...

Chenmin Wu, H. Jia, Yang Liu et al. · 0 citations

Prompts Don't Protect: Architectural Enforcement via MCP Proxy for LLM Tool Access Control

This work identifies a critical gap: when unauthorized tools are visible in an agent's context, models select them in 48-68% of adversarial scenarios, even when explicitly instructed not to, and proposes a proxy-enforced attribute-based access control layer for MCP that filters tool registries at discovery time.

Rohith Uppala · 1 citation

SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces

This work presents SkillSafetyBench, a runnable benchmark for evaluating skill-facing safety failures, and suggests that agent safety depends not only on model-level alignment, but also on how agents interpret skills, trust workflow context, and act through executable environments.

Chang Jin, Anr'an W'ang, Zeming Wei et al. · 11 citations · ⚡1

The Autonomy Tax: Defense Training Breaks LLM Agents

These findings demonstrate that current defense paradigms optimize for single-turn refusal benchmarks while rendering multi-step agents fundamentally unreliable, necessitating new approaches that preserve tool execution competence under adversarial conditions.

Li Li, Yue Zhao · 8 citations
#artificial intelligence Review Aug 2026

LLM-Based Agents for Software and Systems Security: Approaches, Applications, and Assessment

A systematic literature review of technical approaches, including agent architecture, perception, memory, reasoning and planning, action space, orchestration, and self-improvement, reveals a field that has built agents able to act but not yet agents whose authority is bounded or whose behavior is auditable.

Jing-Jing Nie, Jiawei Guo, Krishna Meda et al. · 0 citations
#artificial intelligence Review Aug 2026

LongPIBench: A Long-Context Benchmark for Prompt Injection

LongPIBench is introduced, a long-context benchmark for prompt injection covering 4 realistic application scenarios: paper peer review, resume screening, code review, and email summary, and the evaluation results reveal significant vulnerabilities of prompt injection defenses under long-context settings.

Yupei Liu, Yuqi Jia, N. Gong et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Layered LLM Defenses as an Ensemble: Access Tiers, Inference Cost, and the Measured Failure Correlation Between Defense Layers

How a stack behaves, and the quantities a defender cares about diverge: coverage saturates within a tier, cost rises by class, false refusals accumulate as a union, and residual attack success falls multiplicatively only under independence.

Abrar Alotaibi, M. S. Jabbar, Sadam Al-Azani et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Compared to What? A Human-Anchored Security Benchmark for LLM-Generated Infrastructure-as-Code

GenIaC-SecBench is introduced, a benchmark of 100 deployment scenarios stratified by architectural complexity, evaluated across 12 model configurations from four vendors, producing 1,196 IaC artifacts scanned by three independent policy engines (Checkov, Trivy, KICS).

Animesh Shaw · 1 citation

From tech blogs

See all →
Google DeepMind Blog Jul 17, 2026

Introducing Gemini 3.5 Flash Cyber

Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.