Synthetic data releases are increasingly proposed in the literature as a means of sharing realistic data replicas in lieu of sensitive private datasets. Even when the worst-case privacy leakage of such releases is bounded by means of differential privacy (DP), in practice a residual risk remains. Membership inference a...
Yi-Dan Sun, Viktor Schlegel, Srinivasan Nandakumar et al.· 0 citations
CS-Guard, the first benchmark to systematically evaluate guardrails for code generation security, finds that current guardrails perform poorly against malicious code-generation re- quests, and uses a modular three-layer guardrail taxonomy that lets developers register guardrails for evaluation.
Directional ablation removes an aligned language model's ability to refuse by projecting a single"refusal direction"out of the weights that write the residual stream. It needs no gradient-based training and no optimization, only a few hundred contrastive prompts, which makes it the canonical white-box attack on open-we...
Federated learning shares model updates rather than raw data, yet these updates can be inverted to reconstruct the clients'training data. Analytic reconstruction attacks, which invert a gradient in closed form, degrade as the batch grows: prior single-round attacks recover only about half of a batch of size $100$ even...
Saeed Shariati, M. A. Meybodi· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
On August 2, 2026, the obligations of Article 50 of the EU AI Act took effect, requiring generative AI providers to mark the content their systems produce and ensure it can be detected as AI-generated. Days later, Anthropic disclosed that every Claude model released after that date embeds a watermark based on SynthID-T...
Alexander Nemecek, Vipin Chaudhary, Erman Ayday· 0 citations
Large language model safety and security research is preoccupied with, among other things, detecting and preventing jailbreak attacks: alignment bypasses that allow an adversarial user to elicit unwanted or harmful outputs from models. Arbitrary cipher, or covert communication, attacks are one such type of jailbreak an...
A secure adaptive framework for authentication in multi-zone networks (SAFA-MZ), a causal meta-learning framework for distributed PLA (DPLA) in NTNs, and a model-agnostic meta-learning strategy with invariant risk minimization (IRM) and causal consistency regularization for fast adaptation to unseen NTN environments wi...
Parsa Rajabi, M. Abedi, N. Mokari et al.· 0 citations
This work presents MMPIBench, a reproducible benchmark that measures how far each injected instruction travels through the agent, from perception through planning to the tool call, and extends the benchmark to audio, the only other raw perceptual channel current frontier models accept.
Retrieval-augmented generation (RAG) grounds a language model in retrieved documents, which reduces hallucination but creates a new attack surface: if retrieved text is tampered with, the model may repeat the falsehood. We study how much a small quantized model, Llama 3.1 8B, degrades when a fraction of its retrieved c...
A conceptual defensive architecture separates evidence analysis, adjudication, and operational authority and connects coverage to fallible evidence checking, proper scoring of factual grounding, complete resource accounting, service capacity, and a response model that includes mitigation delay.
This paper presents an end-to-end evaluation framework for image-triggered command injection against computer-use agents (CUAs). The goal is to test whether a local visual patch can induce verifiable environmental consequences along the full chain of screenshot input, VLM generation, action parsing, and environment exe...
Zhihao Liu, Hongyu Sun, Zhiyuan Fu et al.· 0 citations
Large language model (LLM) security has largely focused on role-playing jailbreaks, with less attention to what happens when a user asks an LLM to verify an identity claim through a test designed by the model itself. We study this behavior through a staged developer-identity experiment with ChatGPT, Claude, Qwen, Mistr...