Skip to content
Preprint

ROSETTA: Efficient and Accurate Privacy-Preserving LLM Decoding via Hybrid CKKS/TFHE Evaluation

Sep 2026 · 0 citations · 64 references
Computer Science

TL;DR

ROSETTA is proposed, a hybrid CKKS/TFHE framework that overcomes inefficiency in evaluating nonlinear operations, which incur substantial overhead and dominate the decode stage and achieves up to $4.8\times$ Softmax speedup and $1.5$--$2.1\times$ end-to-end speedup over the SOTA framework CacheMir.

Abstract

Generative large language models (LLMs) have achieved state-of-the-art performance on many real-world tasks such as code generation and question answering. These models predominantly rely on an autoregressive decoding strategy that generates output tokens sequentially. However, their pervasive deployment raises serious privacy concerns, motivating private inference frameworks based on fully homomorphic encryption (FHE). A major limitation of existing FHE frameworks is their inefficiency in evaluating nonlinear operations, which incur substantial overhead and dominate the decode stage. In this paper, we propose ROSETTA, a hybrid CKKS/TFHE framework that overcomes this limitation. We first observe that nonlinear operations in the decode stage exhibit heterogeneous workload patterns, which can be handled effectively via a hybrid approach. We then realize this with two key contributions: 1) an adaptive segmented lookup-table protocol based on TFHE that enables efficient and accurate evaluation of nonlinear operations; and 2) a scheme-aware operator-selection framework that automatically assigns each nonlinear operator to CKKS or TFHE to minimize end-to-end decoding latency. We demonstrate that ROSETTA achieves up to $4.8\times$ Softmax speedup and $1.5$--$2.1\times$ end-to-end speedup over the SOTA framework CacheMir.

View source

Similar papers

Preprint Sep 2026

An Open-Source End-to-End FHE Implementation for Privacy-Preserving Llama 3 8B Inference

Odin is the first open-source end-to-end GPU CKKS implementation of Llama-3, an FHE inference system that co-designs ciphertext packing and model execution for Llama and uses a feature-major cross-layer layout to unify residual connections and layer interfaces.

Yu-Hang Fan, Yu-Si Chen, Kan-Yu Ye et al. · 0 citations
Preprint Aug 2026

SPADE: Speculative Decoding for Precise and Low Cost Distributed Edge Cloud Inference

This work presents a distributed inference framework that integrates speculative decoding across edge and cloud, and shifts the bulk of computation to the edge, significantly lowers inference time and cloud cost, and preserves the accuracy of the big model without any retraining requirement.

D. J. Bajpai, K. Upadhyay, M. Hanawal · 0 citations
2026

PI-SAFE: Practical Privacy-Preserving LLM Inference With Adversarial Fine-Tuning for Optimized Utility

Cloud-based Large Language Model (LLM) inference services typically require users to submit plain-text inputs, thereby posing severe privacy risks. Existing privacy-preserving paradigms are mostly task-specific and often necessitate pervasive modifications to the entire server-side model. This reliance introduces subst...

Wentao Zhong, Yu-Ting Li, Di-Cong Yu et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Auditing and Mitigating Privacy Leakage in Cloud-Edge Collaborative Decoding

CoVeil is proposed, a defense mechanism which dynamically optimizes transmitted signals to suppress leakage during decoding time while preserving the collaborative quality, and consistently improves the privacy-utility trade-off over existing baselines by reducing data leakage.

Ke-Jia Zhang, Tianyuan Zou, Zi-Xuan Gu et al. · 0 citations
Open access Sep 2026

SAFEDT: A Privacy-Preserving Decision Tree Inference Framework Using Homomorphic Encryption

With the increasing application of machine learning in sensitive domains such as healthcare, the demand for privacy-preserving techniques in decision tree inference has grown urgent. Traditional decision trees, when executed in untrusted environments, pose a risk of exposing sensitive data to untrusted parties. Therefo...

Shi-Wen Wei, Zhi-Li Chen, Yan-Zhuo Yu et al. · 0 citations
Preprint Sep 2026

Memory-Efficient Designs for Word-Wise Universal Fully Homomorphic Encryption

Fully Homomorphic Encryption (FHE) enables computation on encrypted data, preserving privacy throughout analysis. While its privacy is very strong, FHE is much slower to execute than the original computation. In particular, due to the recent success in accelerating its compute, the performance bottleneck shifts to the...

A. W. B. Yudha, Erwin Eko Wahyudi, R. Rajagede et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.