Skip to content
Preprint

Understanding the Security Boundary of Obfuscation-based On-Device LLM Protection

Sep 2026 · 0 citations · 55 references
Computer Science

TL;DR

This paper formalizes a set of obfuscation primitives, defined as dual-tuples of linear computations satisfying specific algebraic properties, and demonstrates that the matrix-level weight transformations of several representative efficient TSLP frameworks can be expressed as compositions of these primitives.

Abstract

Trusted Execution Environments (TEEs) offer a promising mechanism for safeguarding the intellectual property of on-device Large Language Models (LLMs). To overcome the inherent computational bottlenecks of TEEs, existing TEE-Shielded LLM Partition (TSLP) methods apply efficient obfuscation schemes to computationally intensive layers, offloading them to external GPUs while retaining only lightweight operations within the TEE. Although a growing body of TSLP-based approaches has emerged, these defense mechanisms remain largely heuristic. Consequently, some methods are proven vulnerable to certain specialized adversarial attacks designed to exploit their specific architectural implementations. To overcome the limitations of these heuristic designs, this paper addresses a fundamental research question: can we establish common primitives to unify representative prior methodologies, characterize the security boundary of their compositions, and systematically extend them? To this end, we formalize a set of obfuscation primitives, defined as dual-tuples of linear computations satisfying specific algebraic properties. We demonstrate that the matrix-level weight transformations of several representative efficient TSLP frameworks can be expressed as compositions of these primitives; consequently, the canonical form of these primitive compositions, denoted as \priorboundary, defines the security boundary of this primitive family. We then expose the vulnerabilities of \priorboundary through a novel primitive-guided attack methodology, \sysattack, demonstrating a shared vulnerability in several prominent TSLP methods published in top-tier venues, such as ArrowCloak (Security'25), TSQP (S\&P'25), and LoRO (NeurIPS'25). Finally, we introduce two novel obfuscation primitives and integrate them with existing constructs to formulate \sysdefense, extending the prior security boundary \priorboundary.

View source

Similar papers

Preprint Sep 2026

Towards TEE-Certified DP: Verifiable Differentially Private Training on Legacy GPUs

Wide adoption of machine learning has created growing policy and regulatory demand for protecting sensitive training data, with differential privacy (DP) emerging as a key mechanism. Yet a less-studied problem is how to certify the faithful execution of DP during training: an external verifier should be able to check t...

Li Ge, Wenjie Qu, Weitao Feng et al. · 0 citations
Open access Aug 2026

Towards Semantic-Preserving Obfuscation for Analysis-Resistant EVM Bytecode

With the transparency of the Ethereum platform, deployed smart contracts remain permanently public, exposing their virtual machine code to risks such as reverse engineering, control-flow analysis and malicious behavior identification. Although several obfuscation approaches for the EVM have been proposed, existing solu...

Dai Dinh Nguyen, Lại Minh Tuấn · 0 citations
Preprint Sep 2026

Type-Directed, Secure-by-Construction Enclave Partitioning for LLVM

Trusted Execution Environments (TEEs) provide hardware-supported isolation through enclaves that protect code and data independently of software abstractions. However, TEEs alone cannot enforce information-flow security. This problem is further aggravated in LLVM-like low-level languages that allow unrestricted pointer...

Wesley B. Nuzzo, Samuel Dodson, Benjamin Houle et al. · 0 citations
Preprint Aug 2026

Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure

Design guidance for measurement under strategic optimization is distill design guidance for measurement under strategic optimization: held-out probes retain validity only on non-enumerable axes; gates must measure held-out performance, not just correctness; and a transfer rate is interpretable only with per-failure mec...

Víctor Gallego · 0 citations
#artificial intelligence Preprint Sep 2026

Permutation-Based Stegomalware in Large Language Models: Threats and Countermeasures

This paper demonstrates the full potential of behavior-preserving symmetries as a defense against stegomalware, as well as the risks these symmetries pose when exploited by attackers, and quantifies the loss in model performance associated with applying these methods.

Danny Wood, James Stringer · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.