2026· IEEE Transactions on Information Forensics and Security· Vol 21, pp. 8270-8285· 0 citations· 54 references
Abstract
Cloud-based Large Language Model (LLM) inference services typically require users to submit plain-text inputs, thereby posing severe privacy risks. Existing privacy-preserving paradigms are mostly task-specific and often necessitate pervasive modifications to the entire server-side model. This reliance introduces substantial computational overhead, elevated deployment costs, and an inevitable privacy-utility trade-off. In this paper, we propose <monospace>PI-SAFE</monospace>, a novel framework enabling privacy-preserving LLM inference for both text generation and classification tasks, notably without requiring any alterations to the server-side model. By partitioning a pre-trained LLM between the client and the server, <monospace>PI-SAFE</monospace> facilitates collaborative inference via the transmission of obfuscated intermediate representations rather than raw text. To defend against reconstruction attacks, we introduce a Dual-constraint Adversarial Fine-Tuning (AdvFT) mechanism, which empowers the client to reshape the output feature distribution by covertly inserting adapter modules. To further counter powerful adversaries equipped with massive prior data who employ Deep Neural Networks (DNNs) for inverse fitting, we propose an enhanced framework, <monospace>PI-SAFE</monospace><inline-formula> <tex-math notation="LaTeX">${}^{\boldsymbol {+}}$ </tex-math></inline-formula>. This framework introduces a random prefix-based feature obfuscation mechanism. By employing a locally generated, fixed secret token sequence for feature fusion, <monospace>PI-SAFE</monospace><inline-formula> <tex-math notation="LaTeX">${}^{\boldsymbol {+}}$ </tex-math></inline-formula> fundamentally disrupts the topological isomorphism of the feature mapping space. Extensive experiments demonstrate that, under equivalent privacy guarantees, <monospace>PI-SAFE</monospace><inline-formula> <tex-math notation="LaTeX">${}^{\boldsymbol {+}}$ </tex-math></inline-formula> significantly outperforms existing baselines in inference utility. Crucially, it reduces the success rate of optimization-based reconstruction attacks to near zero and degrades attribute inference to random-guessing levels, exhibiting exceptional empirical defense and computational efficiency.
Cloud-based large language models (LLMs) require users to disclose plaintext data to service providers, creating privacy risks in sensitive domains. Existing privacy-preserving approaches often trade utility for protection, incur substantial computational or communication overhead, remain vulnerable to reconstruction f...
Jeongho Yoon, Chanhee Park, Yong-Chan Chun et al.· 0 citations
Gecko is presented, designed to limit this additional risk while retaining a compact encrypted predictor, and formalizes ideal independence and information-preservation conditions as design guidance, then separately evaluate component-reuse extraction attacks.
Cheng'an Wei, Kai Chen, Yue Zhao et al.· 0 citations
CoVeil is proposed, a defense mechanism which dynamically optimizes transmitted signals to suppress leakage during decoding time while preserving the collaborative quality, and consistently improves the privacy-utility trade-off over existing baselines by reducing data leakage.
Ke-Jia Zhang, Tianyuan Zou, Zi-Xuan Gu et al.· 0 citations
P2Skill is proposed, a prompt-based skill distillation method in which a local small language model (SLM) autonomously performs decomposition, PII-aware routing, paraphrasing, and reconstruction by following the skill prompts.
M. Ryu, Geunpyo Park, Sungjoon Lee et al.· 1 citation· ⚡1
A privacy-preserving zk-SNARK-based audit framework that searches for probes designed in the spirit of adversarial examples to amplify logit drift between an approved model and a modified deployment and demonstrates that token-based probes consistently deliver the strongest mean sensitivity across models and GPU platfo...
Cameron Wilding, Mina Shaker, Fatemeh Ganji· 0 citations
This work addresses leakage through a learned obfuscate-and-recover scheme that protects participants' private datasets while still allowing an independently deployable model to be trained on the server side, making split-based federated LLM fine-tuning practically viable.
Heng Jin, Chao-Yu Zhang, He-Xuan Yu et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.