Skip to content

Securing LLMs in the Wild: Privacy and Security Challenges at the Edge

Jul 2026 · arXiv.org · Vol abs/2607.13088 · 0 citations · 58 references
Computer Science

TL;DR

A deployment-centric taxonomy organized around three architectural constraints is introduced, derive a unified constraint model that quantifies when unsafe optimizations become unavoidable, linking each wall to specific attack surfaces, and proposes the Secure Operational Efficiency Score (SOES), a holistic metric balancing task accuracy, jailbreak resistance, and privacy against energy, memory, and latency.

Abstract

Large Language Models (LLMs) are rapidly moving from research settings into the wild, deployed on enterprise infrastructure, personal devices, and edge platforms. While cloud deployments offer scalable compute, concerns over data sovereignty, compliance, latency, and third-party dependence are driving organizations toward edge and on-premise LLMs. This shift introduces new security and privacy challenges: limited compute and memory force aggressive optimizations, including quantization, pruning, model partitioning, and parameter-efficient adaptation, each of which can introduce vulnerabilities and reshape the threat landscape. We describe this tension as the Security-Efficiency Paradox, mechanisms that improve efficiency may weaken robustness, expose new attack surfaces, or increase privacy risks. We examine how compression can degrade safety alignment, how partitioned inference enables reconstruction attacks, and how continuous local adaptation may cause privacy leakage and model drift. To analyze these risks, we introduce a deployment-centric taxonomy organized around three architectural constraints: the Memory Wall, the Quadratic Wall, and the Compute Wall. We derive a unified constraint model that quantifies when unsafe optimizations become unavoidable, linking each wall to specific attack surfaces. Building on this model, we propose the Secure Operational Efficiency Score (SOES), a holistic metric balancing task accuracy, jailbreak resistance, and privacy against energy, memory, and latency, enabling practitioners to configure edge LLMs under real-world hardware limits. We further present a practical decision procedure and targeted mitigations for each optimization-induced vulnerability. Together, these contributions provide a co-designed framework for jointly evaluating security, privacy, and efficiency, laying a foundation for securing edge-native intelligent systems.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

Auditing and Mitigating Privacy Leakage in Cloud-Edge Collaborative Decoding

CoVeil is proposed, a defense mechanism which dynamically optimizes transmitted signals to suppress leakage during decoding time while preserving the collaborative quality, and consistently improves the privacy-utility trade-off over existing baselines by reducing data leakage.

Ke-Jia Zhang, Tianyuan Zou, Zi-Xuan Gu et al. · 0 citations
Open access Aug 2026

A Privacy-Preserving Middleware Architecture for Detecting Prompt Injection and Sensitive Data Exposure in Large-Language-Model Interactions

A privacy-preserving hybrid middleware architecture that enforces a local trust boundary as its primary design constraint that is model-agnostic, requires no retraining of the underlying LLM, and is compatible with black-box API deployments is proposed and evaluated.

Adam Ait Hsine, A. Arabo · 0 citations
Book Open access Aug 2026

The 2nd SeT-LLM Workshop on Secure and Trustworthy Large Language Models

Large language models (LLMs) are increasingly embedded as core components of data-centric systems, supporting analytical decision making, and automated reasoning over large-scale, heterogeneous datasets. Yet their deployment in open-world environments raises fundamental challenges to security and trustworthiness: LLMs...

Lu Lin, Jinghui Chen, Ting Wang et al. · 0 citations
Review Open access 2026

Privacy and Security in Knowledge Distillation for Federated Learning: A Survey

Knowledge distillation (KD) is increasingly used in federated learning (FL) because it enables clients to exchange predictions, features, prototypes, or synthetic knowledge rather than full model parameters. This change can reduce communication and support heterogeneous models, but it also changes the privacy and secur...

Hamza Reguieg, Essaid Sabir, M. El Kamili · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.