Skip to content

Demystifying the Privacy-Utility Trade-off in LLM Interactions

Sep 2026 · 1 citation · 37 references
Computer Science

TL;DR

By distilling a lightweight model Veilmind-4B to drive a dynamic extraction-sanitization-restoration pipeline, this approach reaches a low-leakage privacy point while preserving substantially higher response utility than existing privacy-oriented baselines, advancing the privacy-utility trade-off toward the Pareto frontier.

Abstract

The integration of Large Language Models into daily tasks relies on context-rich instructions, inevitably exposing sensitive user information. Current privacy-preserving methods typically employ context-agnostic static rules, causing severe utility degradation. However, the specific mechanisms governing how sanitization impacts downstream performance remain largely underexplored. To address this, we conduct a systematic analysis to deconstruct the privacy-utility trade-off, uncovering three underlying mechanisms: (1) Context-Dependent Utility, which first establishes when to sanitize by revealing that data value shifts from critical constraints to dispensable noise based on user intent; (2) Strategic Adaptation, which subsequently determines how to sanitize by dictating that the choice between removal and replacement depends on the task's reliance on factual integrity versus structural coherence; and (3) Combinatorial Interplay, which finally extends the protection scope by demonstrating that attributes form a semantic web of synergistic dependencies or antagonistic redundancies. Guided by these insights, we introduce an intent-driven local protection framework. By distilling a lightweight model Veilmind-4B to drive a dynamic extraction-sanitization-restoration pipeline, our approach reaches a low-leakage privacy point while preserving substantially higher response utility than existing privacy-oriented baselines, advancing the privacy-utility trade-off toward the Pareto frontier.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

LLMs as Linguistic Chameleons: Decoupling Semantics and Structure for Privacy-Preserving Communication

As Large Language Model (LLM) APIs become increasingly integrated into privacy-sensitive workflows, ensuring inference-time privacy without compromising task utility remains a major challenge. Existing approaches preserve most of the original semantic content to maintain downstream performance, but this also leaves exp...

Yu-Zhu Mao, Liang Zhao · 0 citations
Preprint Aug 2026

Beyond Direct Identifiers: Probabilistic Privacy Risk Estimation for Privacy-Conscious LLM Query Delegation

This work investigates a probabilistic variant of PCD, where an LLM-driven probabilistic estimation of k-anonymity is augmented with an LLM-driven probabilistic estimation of k-anonymity, and proposes k-anonymity as a useful auxiliary metric for tackling PCD.

Si-Yan Li, Yu Zhou, Julia Hirschberg · 2 citations
2026

PI-SAFE: Practical Privacy-Preserving LLM Inference With Adversarial Fine-Tuning for Optimized Utility

Cloud-based Large Language Model (LLM) inference services typically require users to submit plain-text inputs, thereby posing severe privacy risks. Existing privacy-preserving paradigms are mostly task-specific and often necessitate pervasive modifications to the entire server-side model. This reliance introduces subst...

Wentao Zhong, Yu-Ting Li, Di-Cong Yu et al. · 0 citations
Preprint Sep 2026

Decoding the Legalese: A Scalable and Quantitative Framework for Analyzing Corporate Privacy Policies

Even though privacy policies are the primary mechanism organizations use to disclose how they collect, process, and share personal data, they are difficult for average users to interpret, perhaps by design, due to their verbosity and dense legal language. Importantly, there is a lack of standardized metrics that charac...

Jia-Ming Tang, Chen-Lan Wang, Ming-Yan Liu et al. · 1 citation · ⚡1
#machine learning Preprint Oct 2026

Intent-Hiding Jailbreaks: An Information-Theoretic Framework for Compositional Attacks

Recent work has shown that large language models (LLMs) can be vulnerable to jailbreak attacks in which harmful intent is obscured through composition with benign tasks. A harmful request refused in isolation may elicit a different response when embedded within a larger, seemingly benign query. We study these compositi...

Feng-Wei Tian, Ravi Tandon · 0 citations
Preprint Sep 2026

PIMENTO: A Privacy Framework for Querying Text

Currently, there are two state-of-the-art, complementary privacy guarantees: contextual integrity (CI) for what may flow, and differential privacy (DP) for what may be inferred. Yet neither maps cleanly onto natural language, leaving existing approaches unable to provide these guarantees for analytics over unstructured...

Mushtari Sadia, Ang Chen, Amrita Roy Chowdhury · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.