Skip to content

Author

Nils Lukas

We have 7 of 46 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Oct 2026

Certification of Real Images through Calibrated Content Authentication

Generative models can synthesize high-quality inauthentic multimedia content that is already being misused at scale. We evaluate twenty deepfake detectors against ten generators released in the last four years and find accuracy decreasing over time, from near-perfect 99.5% to 76%. Adversarial perturbations further redu...

Sarim Hashmi, Abdelrahman W. A. Elsayed, Mohammed Talha Alam et al. · 0 citations
#natural language process... Preprint Oct 2026

Noise Out, Bias In: Targeted Bias Injection in Diffusion Language Models via Closed-Loop Activation Steering

Masked diffusion language models (dLLMs) generate text by iteratively denoising masked positions, re-predicting each token multiple times before it is committed. An autoregressive decoder exposes an answer's distribution once, at the step that commits it; a dLLM exposes it at every denoising step before commitment, and...

Sarim Hashmi, Mukul Ranjan, Abdelrahman W. A. Elsayed et al. · 0 citations

Watermarking Should Be Treated as a Monitoring Primitive

It is shown that even zero-bit watermarking supports internal attribution under per-entity multi-key deployments without explicitly encoding identity, and external identification in selected text and image configurations is demonstrated.

Toluwani Aremu, Nils Lukas, Jie Zhang · 1 citation
Conference Open access Sep 2026

Addressing Overcommitment in the Reasoning of Gendered Economic Memes Under Multimodal Ambiguity

Multimodal meme understanding is increasingly used to analyze socially sensitive content, yet existing models often exhibit biased behavior when interpreting economic dependence and social roles under ambiguity. Many memes express economic relationships through sparse text or symbolic visual cues, providing insufficien...

Kushal Kanwar, Dushyant Singh Chauhan, Kapil Rana et al. · 0 citations
#machine learning Preprint Aug 2026

Context Inference Attacks Without Jailbreaks

This work introduces and formalizes context-inference attacks through a security game and evaluates three settings under decreasing attacker knowledge and increasingly indirect delivery of the context: a known context, an unknown context, and a context the agent retrieves through its own tool calls.

Prince Jha, Samuele Poppi, Nils Lukas · 0 citations
#machine learning Preprint Sep 2026

CopyShield: A Cross-Level Benchmark of Copyright Defenses in LLMs

Large language models can reproduce memorized text verbatim, yet copyright defenses are usually evaluated under incompatible protocols. We introduce CopyShield, a controlled benchmark comparing three representative defenses at distinct intervention levels: contrastive decoding (output), Direct Preference Optimization (...

Maryam Alshehyari, Dushyant Singh Chauhan, Samuele Poppi et al. · 0 citations

A Gravitational Interpretation of Fine-Tuning Reversion

Fine-tuning on harmless data can partially undo behaviors acquired earlier in training. Safety can erode under benign post-alignment updates, unlearned capabilities can re-emerge, latent traits can transfer through apparently unrelated supervision, and related post-alignment fragility appears in other generative settin...

Samuele Poppi, Nils Lukas · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.