Skip to content
Preprint

An Open-Source End-to-End FHE Implementation for Privacy-Preserving Llama 3 8B Inference

Sep 2026 · 0 citations · 15 references
Computer Science

TL;DR

Odin is the first open-source end-to-end GPU CKKS implementation of Llama-3, an FHE inference system that co-designs ciphertext packing and model execution for Llama and uses a feature-major cross-layer layout to unify residual connections and layer interfaces.

Abstract

Cloud LLM services typically require users to send prompts to a model provider, creating a privacy risk. Fully homomorphic encryption (FHE) lets a server perform inference without decrypting the input, but representing data as ciphertexts adds storage and computational overhead. In CKKS-based LLM inference, the packing scheme maps logical tensors to ciphertexts and slots. It therefore determines the ciphertext count and the homomorphic cost of linear layers, and it constrains how data pass between linear layers, attention, and nonlinear computation. As models and sequences grow, inefficient layouts accumulate encoding, compute, and layout-conversion overhead. We present Odin, an FHE inference system that co-designs ciphertext packing and model execution for Llama. Starting from a THOR-style baseline whose bottleneck is weight encoding, Odin uses a feature-major cross-layer layout to unify residual connections and layer interfaces, and builds transient intra-operator layouts for linear projections and attention. This reduces redundant plaintext encoding of weights in wide projections. Within attention, QK^T produces scores that Softmax can consume directly, and PV consumes the resulting probabilities, avoiding intermediate repacking. For nonlinear ops, we use minimax polynomial approximation with input-range control and joint error allocation guided by model quality, reducing polynomial degree and multiplicative depth. To our knowledge, Odin is the first open-source end-to-end GPU CKKS implementation of Llama-3. With Llama-3-8B weights and a 128-token input, Odin evaluates all 32 Transformer layers on a single NVIDIA H100 80 GB GPU. Server-side end-to-end FHE evaluation takes 366.4 s and 58.9 GiB peak device memory. Under the same model, input, CKKS parameters, and hardware, THOR takes 1651.9 s, a 4.51x speedup.

View source

Similar papers

Preprint Aug 2026

Gecko: Fast Private Inference via Secure Public Encoder Offloading

Gecko is presented, designed to limit this additional risk while retaining a compact encrypted predictor, and formalizes ideal independence and information-preservation conditions as design guidance, then separately evaluate component-reuse extraction attacks.

Cheng'an Wei, Kai Chen, Yue Zhao et al. · 0 citations
Open access Aug 2026

EHEIR: Efficient Homomorphic Encrypted Inference via Architectural Redesign

This work presents a framework that reformulates HE-aware model design as a constrained neural architecture search problem, where the objective is to identify architectures that are both cryptographically feasible and computationally efficient while preserving task performance.

Reeshav Chowdhury, Anoop Mishra, Deepak Khazanchi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

HEAT: Faster Fully Homomorphic Inference via Approximations-Weights Co-Adaptation

Fully homomorphic encryption (FHE) allows a server to run a language model directly on encrypted user prompts, but current approaches remain prohibitively slow. Ciphertexts natively support only addition, multiplication, and rotation, and multiplications may be composed only to a bounded depth before a costly bootstrap...

Alessandro Zirilli, Davide Marincione, Evgenios M. Kornaropoulos et al. · 0 citations
Preprint Sep 2026

SpliTEE: Fast and Private LLM Inference by Coupling GPU-Assisted Trusted Execution Environments with Differential Privacy

User prompts provided to large language models (LLMs) may contain sensitive or private information that can be misused by remotely deployed models, such as through inadvertent memorization during retraining. One way to protect user prompts is to execute the LLM inside a trusted execution environment (TEE), with the gua...

Shashie Dilhara Batan Arachchige, Robin Carpentier, H. Asghar et al. · 0 citations
Sep 2026

Shuffling is Not Enough: Breaking Permutation-Based Model Confidentiality in Hybrid FHE Inference

It is shown that input DP is orthogonal to model confidentiality and that the local-DP premise required for shuffle amplification cannot hold under correctness-bounded noise, and that the leaked spectra enable fingerprinting, lineage attribution, and improved logit-based extraction, while suppressing them destroys infe...

Jiseung Kim, Hyung Tae Lee · 0 citations
Preprint Sep 2026

Memory-Efficient Designs for Word-Wise Universal Fully Homomorphic Encryption

Fully Homomorphic Encryption (FHE) enables computation on encrypted data, preserving privacy throughout analysis. While its privacy is very strong, FHE is much slower to execute than the original computation. In particular, due to the recent success in accelerating its compute, the performance bottleneck shifts to the...

A. W. B. Yudha, Erwin Eko Wahyudi, R. Rajagede et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.