Skip to content

PrivDrift: Auditing User-Secret Leakage Under Topic Drift in Active LLM Conversations

Sep 2026 · 0 citations · 17 references
Computer Science

TL;DR

PrivDrift is introduced, a benchmark for auditing whether user-disclosed secrets remain recoverable after conversational topic drift and persuasion-based probing, suggesting that privacy risk in active LLM contexts should be evaluated as a persistent behavioral failure mode rather than only as training-data memorization or immediate jailbreak behavior.

Abstract

Large language models increasingly operate as persistent assistants in user-facing, shared-session, and tool-augmented settings. When users disclose sensitive information during an active conversation, that information may remain behaviorally recoverable through later prompts even after the dialogue shifts to unrelated topics. We introduce PrivDrift, a benchmark for auditing whether user-disclosed secrets remain recoverable after conversational topic drift and persuasion-based probing. PrivDrift contains 1,000 controlled multi-turn dialogues with seeded secrets, content-dense drift turns, and standardized extraction probes. Across three LLMs with extended context windows, dialogue-level hybrid leakage remains substantial, ranging from 38.7% to 54.6%, and varies strongly by model, secret type, and persuasion intensity. Within the tested drift window, additional topic drift does not reliably reduce leakage, suggesting that privacy risk in active LLM contexts should be evaluated as a persistent behavioral failure mode rather than only as training-data memorization or immediate jailbreak behavior.

View source

Similar papers

Preprint Sep 2026

Can Prompt Anonymity Protect Your Identity From LLM Providers?

User conversations with large language models (LLMs) often contain highly sensitive personal information that can be exploited by LLM providers to create detailed user dossiers, enable targeted advertising, and train more powerful models. To protect user privacy, anonymizing LLM proxies have emerged as a practical solu...

Dzung Pham, Dillon Sheils, Naina Singh et al. · 0 citations
Preprint Sep 2026

Your Model Is Leaking: Covert Information Transfer through LLM Residual Streams

Privacy-sensitive organizations may run large language models (LLMs) in restricted or air-gapped environments while exporting selected diagnostic artifacts. We show that a compromised runtime component can hide sensitive information in intermediate activations that are allowed to leave the restricted environment. An of...

Ming-Yuan Li, Yan-Na Jiang, Guang-Sheng Yu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Confuse the Model, Control the Flow: Understanding and Mitigating Privacy Leakage from LLM Agents with Information Flow Control

Personal AI agents built on large language models (LLMs) are increasingly given access to a user's private data and communications in order to provide personalized assistance. This access creates a persistent privacy risk: the agent must decide whether a given sensitive information should be disclosed to a particular p...

Minsun Shim, Ramisha Raida Karim, Ruthwik Jakkula et al. · 1 citation
Preprint Aug 2026

Inadvertent Context Leakage in Language Models

Leakage enables two practical attacks: a trained classifier that infers semantic predicates about user memories from routine natural-language outputs, and an RL-trained adversary that extracts full Social Security Numbers from a production-style agent.

Jaiden Fairoze, Neal Mangaokar, Kamalika Chaudhuri et al. · 0 citations
#artificial intelligence Review Sep 2026

ASLEval: Measuring Privacy Exposure Displacement in LLM Agent Sessions

Privacy exposure displacement, the mismatch between a local evaluation proxy and target-grounded session exposure, and ASLEval, an authorization-aware framework that pre-registers a hidden target set, measures all declared visible exits, and reserves internal traces for diagnosis are introduced.

Guo-Xin Wu, Hui-Zhen Huang, Guo-Xiong Long et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Canaries in the Bank: Auditing User-Level Privacy in Private Evolution

A protocol-aware empirical audit is introduced in which the server commits to a single shared candidate bank and replaces roughly 1% of its entries with probes derived from a known, non-private canary, to quantify the gap between formal worst-case privacy and leakage achievable through protocol-valid candidate-bank man...

Sai Aparna Aketi, Enayat Ullah, Shripad Gade · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.