Generative models can synthesize high-quality inauthentic multimedia content that is already being misused at scale. We evaluate twenty deepfake detectors against ten generators released in the last four years and find accuracy decreasing over time, from near-perfect 99.5% to 76%. Adversarial perturbations further redu...
Sarim Hashmi, Abdelrahman W. A. Elsayed, Mohammed Talha Alam et al.· 0 citations
Masked diffusion language models (dLLMs) generate text by iteratively denoising masked positions, re-predicting each token multiple times before it is committed. An autoregressive decoder exposes an answer's distribution once, at the step that commits it; a dLLM exposes it at every denoising step before commitment, and...
Sarim Hashmi, Mukul Ranjan, Abdelrahman W. A. Elsayed et al.· 0 citations
It is shown that even zero-bit watermarking supports internal attribution under per-entity multi-key deployments without explicitly encoding identity, and external identification in selected text and image configurations is demonstrated.
Toluwani Aremu, Nils Lukas, Jie Zhang· arXiv.org· 1 citation
Multimodal meme understanding is increasingly used to analyze socially sensitive content, yet existing models often exhibit biased behavior when interpreting economic dependence and social roles under ambiguity. Many memes express economic relationships through sparse text or symbolic visual cues, providing insufficien...
Kushal Kanwar, Dushyant Singh Chauhan, Kapil Rana et al.· Proceedings of the Thirty-Fi...· 0 citations
This work introduces and formalizes context-inference attacks through a security game and evaluates three settings under decreasing attacker knowledge and increasingly indirect delivery of the context: a known context, an unknown context, and a context the agent retrieves through its own tool calls.
Prince Jha, Samuele Poppi, Nils Lukas· 0 citations
Large language models can reproduce memorized text verbatim, yet copyright defenses are usually evaluated under incompatible protocols. We introduce CopyShield, a controlled benchmark comparing three representative defenses at distinct intervention levels: contrastive decoding (output), Direct Preference Optimization (...
Fine-tuning on harmless data can partially undo behaviors acquired earlier in training. Safety can erode under benign post-alignment updates, unlearned capabilities can re-emerge, latent traits can transfer through apparently unrelated supervision, and related post-alignment fragility appears in other generative settin...
Samuele Poppi, Nils Lukas· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.