Generative models can synthesize high-quality inauthentic multimedia content that is already being misused at scale. We evaluate twenty deepfake detectors against ten generators released in the last four years and find accuracy decreasing over time, from near-perfect 99.5% to 76%. Adversarial perturbations further redu...
Sarim Hashmi, Abdelrahman W. A. Elsayed, Mohammed Talha Alam et al.· 0 citations
This work introduces and formalizes context-inference attacks through a security game and evaluates three settings under decreasing attacker knowledge and increasingly indirect delivery of the context: a known context, an unknown context, and a context the agent retrieves through its own tool calls.
Prince Jha, Samuele Poppi, Nils Lukas· 0 citations
Large language models can reproduce memorized text verbatim, yet copyright defenses are usually evaluated under incompatible protocols. We introduce CopyShield, a controlled benchmark comparing three representative defenses at distinct intervention levels: contrastive decoding (output), Direct Preference Optimization (...
This work introduces ReACT-CLIP, a response-conditioned test-time defense that separately determines how strongly each input should be corrected and whether defensive intervention is necessary, and quantifies this variation using a prediction-instability score computed by Jensen--Shannon divergence and combines it with...
H. Malik, Toluwani Aremu, Samuele Poppi et al.· 0 citations
Fine-tuning on harmless data can partially undo behaviors acquired earlier in training. Safety can erode under benign post-alignment updates, unlearned capabilities can re-emerge, latent traits can transfer through apparently unrelated supervision, and related post-alignment fragility appears in other generative settin...
Samuele Poppi, Nils Lukas· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.