This work develops GREAT, a novel framework for crafting natural distributional backdoors in RLHF, which targets harmful response generation for a vulnerable user subpopulation featured by semantically violent requests paired with emotionally angry triggers.
Subrat Kishore Dutta, Yuelin Xu, P. Pant et al.· arXiv.org· 0 citations
The exact DP constant is pin down for the two that carry the practical weight, counterfactual memorization and adaptive extraction, and it is shown that they do not control each other.
This work presents Replicant, a deep reinforcement learning framework that learns the realistic task of evasion under a strict label-only black-box threat model and demonstrates that learning the task of evasion not only results in stronger attack performance but provides a better signal for hardening malware detectors...
Shae McFadden, Ilias Tsingenopoulos, Mario D'Onghia et al.· 0 citations
LoopHarness is presented, which restores a persistent, non-decaying safety state at the loop level at the loop level, and gives a complete evaluation protocol on native Agent-SafetyBench tasks with paired clean and attacked episodes, an outer-state attack suite whose decisive evidence exists only across iterations, per...
Chenmin Wu, H. Jia, Yang Liu et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
This work identifies a critical gap: when unauthorized tools are visible in an agent's context, models select them in 48-68% of adversarial scenarios, even when explicitly instructed not to, and proposes a proxy-enforced attribute-based access control layer for MCP that filters tool registries at discovery time.
This work presents SkillSafetyBench, a runnable benchmark for evaluating skill-facing safety failures, and suggests that agent safety depends not only on model-level alignment, but also on how agents interpret skills, trust workflow context, and act through executable environments.
These findings demonstrate that current defense paradigms optimize for single-turn refusal benchmarks while rendering multi-step agents fundamentally unreliable, necessitating new approaches that preserve tool execution competence under adversarial conditions.
A systematic literature review of technical approaches, including agent architecture, perception, memory, reasoning and planning, action space, orchestration, and self-improvement, reveals a field that has built agents able to act but not yet agents whose authority is bounded or whose behavior is auditable.
Jing-Jing Nie, Jiawei Guo, Krishna Meda et al.· 0 citations
LongPIBench is introduced, a long-context benchmark for prompt injection covering 4 realistic application scenarios: paper peer review, resume screening, code review, and email summary, and the evaluation results reveal significant vulnerabilities of prompt injection defenses under long-context settings.
This paper proposes optimal testing strategies which can still recover needed test results even if there are cheaters polluting the results, and determines the optimal testing strategies using a dynamic programming method.
How a stack behaves, and the quantities a defender cares about diverge: coverage saturates within a tier, cost rises by class, false refusals accumulate as a union, and residual attack success falls multiplicatively only under independence.
Abrar Alotaibi, M. S. Jabbar, Sadam Al-Azani et al.· 0 citations
GenIaC-SecBench is introduced, a benchmark of 100 deployment scenarios stratified by architectural complexity, evaluated across 12 model configurations from four vendors, producing 1,196 IaC artifacts scanned by three independent policy engines (Checkov, Trivy, KICS).