PrivDrift is introduced, a benchmark for auditing whether user-disclosed secrets remain recoverable after conversational topic drift and persuasion-based probing, suggesting that privacy risk in active LLM contexts should be evaluated as a persistent behavioral failure mode rather than only as training-data memorization or immediate jailbreak behavior.
Abstract
Large language models increasingly operate as persistent assistants in user-facing, shared-session, and tool-augmented settings. When users disclose sensitive information during an active conversation, that information may remain behaviorally recoverable through later prompts even after the dialogue shifts to unrelated topics. We introduce PrivDrift, a benchmark for auditing whether user-disclosed secrets remain recoverable after conversational topic drift and persuasion-based probing. PrivDrift contains 1,000 controlled multi-turn dialogues with seeded secrets, content-dense drift turns, and standardized extraction probes. Across three LLMs with extended context windows, dialogue-level hybrid leakage remains substantial, ranging from 38.7% to 54.6%, and varies strongly by model, secret type, and persuasion intensity. Within the tested drift window, additional topic drift does not reliably reduce leakage, suggesting that privacy risk in active LLM contexts should be evaluated as a persistent behavioral failure mode rather than only as training-data memorization or immediate jailbreak behavior.
User conversations with large language models (LLMs) often contain highly sensitive personal information that can be exploited by LLM providers to create detailed user dossiers, enable targeted advertising, and train more powerful models. To protect user privacy, anonymizing LLM proxies have emerged as a practical solu...
Dzung Pham, Dillon Sheils, Naina Singh et al.· 0 citations
Privacy-sensitive organizations may run large language models (LLMs) in restricted or air-gapped environments while exporting selected diagnostic artifacts. We show that a compromised runtime component can hide sensitive information in intermediate activations that are allowed to leave the restricted environment. An of...
Ming-Yuan Li, Yan-Na Jiang, Guang-Sheng Yu et al.· 0 citations
Personal AI agents built on large language models (LLMs) are increasingly given access to a user's private data and communications in order to provide personalized assistance. This access creates a persistent privacy risk: the agent must decide whether a given sensitive information should be disclosed to a particular p...
Leakage enables two practical attacks: a trained classifier that infers semantic predicates about user memories from routine natural-language outputs, and an RL-trained adversary that extracts full Social Security Numbers from a production-style agent.
Jaiden Fairoze, Neal Mangaokar, Kamalika Chaudhuri et al.· 0 citations
Privacy exposure displacement, the mismatch between a local evaluation proxy and target-grounded session exposure, and ASLEval, an authorization-aware framework that pre-registers a hidden target set, measures all declared visible exits, and reserves internal traces for diagnosis are introduced.
Guo-Xin Wu, Hui-Zhen Huang, Guo-Xiong Long et al.· 0 citations
A protocol-aware empirical audit is introduced in which the server commits to a single shared candidate bank and replaces roughly 1% of its entries with probes derived from a known, non-private canary, to quantify the gap between formal worst-case privacy and leakage achievable through protocol-valid candidate-bank man...
Sai Aparna Aketi, Enayat Ullah, Shripad Gade· 0 citations