Long-running LLM agents are stateful: beyond the transcript they accrete compressed summaries, plaintext memory, pending tool plans, and, under every serving API, a KV cache. Yet today's"forget"operations delete a plaintext memory record and stop, leaving every artifact derived from the revoked information intact. We f...
Chao Yao, Yang-Bo Wei, Zhen Huang et al.· 1 citation
Adaptive HMAS can provide accuracy cost tradeoff for ransomware analysis while retaining support for heterogeneous and incomplete modalities, as well as reducing average analysis cost and average analysis latency.
Prompt injection is widely recognized as a major security threat to AI agents that interact with untrusted external data, such as websites, documents, and emails. Prior work has shown that, in the text domain, black-box prompt injection can achieve near-perfect attack success rates (ASRs). In the image domain, however,...
Sizhe Chen, Yu-Lin Tsai, Ivan Evtimov et al.· 0 citations
This paper analyzes the implementation of blockchain-based integrity mechanisms in Greek Fiscal Electronic Mechanisms (FEMs) and the central tax information system eSEND. The study examines the cryptographic architecture of fiscal devices, including Electronic Cash Registers, Fiscal Printers, Fiscal Signing Machines, a...
This work identifies the attacker's adaptive search over the system attack surfaces as an important and underexplored security risk for tool-using agents and suggests that agentic security evaluations should characterize both the attacker's search procedure and compute budget.
D. M. Nguyen, Joon Sik Kim, Blazej Manczak et al.· 1 citation
This paper presents a meta-synthesis that draws together four constituent studies covering adversarial machine learning, AI-powered anomaly detection in cloud environments, automated vulnerability patching by multi-agent large language model (LLM) pipelines, and the broader landscape of securing AI systems across their...
Harsh Verma· International Journal of Sci...· 0 citations
DirBucket is the only method that consistently achieves strong target detection with no non-target activation, detecting non-compliance in every audit within 23 audited answers on the authors' primary benchmark, and the results suggest that embedding-space watermarking can make document reuse in third-party RAG statist...
Alexandr Goultiaev Tolstokorov, K. Mouratidis, Javad Dogani et al.· 0 citations
The findings support selfie-capture motion as a low-friction auxiliary signal, while leaving cross-device, cross-session, and real injection-attack evaluation as necessary next steps.
Erkka Rantahalvari, Olli Silvén, Z. Boulkenafet et al.· arXiv.org· 1 citation
Specification gaming under Reinforcement Learning (RL) is known to cause LLMs to develop sycophantic, manipulative, or deceptive behavior, yet the conditions under which this occurs remain unclear. We train 11 instruction-tuned LLMs (0.5B-14B) with on-policy RL across 3 environments and find that model size acts as a s...
Leon Eshuijs, Shihan Wang, Antske Fokkens· 0 citations
We introduce Stateless Bernoulli Watermarking (SBW), a new statistical watermark for Large Language Models that determines green list membership through independent per-token Bernoulli trials. Unlike KGW's vocabulary permutation or SynthID's multi-layer tournament, SBW requires only a single comparison per token agains...
CRAW is introduced, a codec-robust audio watermarking framework that jointly improves robustness against neural re-synthesis while maintaining high perceptual quality and achieves state-of-the-art robustness against neural codecs, denoisers, and vocoders.