ROSETTA is proposed, a hybrid CKKS/TFHE framework that overcomes inefficiency in evaluating nonlinear operations, which incur substantial overhead and dominate the decode stage and achieves up to $4.8\times$ Softmax speedup and $1.5$--$2.1\times$ end-to-end speedup over the SOTA framework CacheMir.
Abstract
Generative large language models (LLMs) have achieved state-of-the-art performance on many real-world tasks such as code generation and question answering. These models predominantly rely on an autoregressive decoding strategy that generates output tokens sequentially. However, their pervasive deployment raises serious privacy concerns, motivating private inference frameworks based on fully homomorphic encryption (FHE). A major limitation of existing FHE frameworks is their inefficiency in evaluating nonlinear operations, which incur substantial overhead and dominate the decode stage. In this paper, we propose ROSETTA, a hybrid CKKS/TFHE framework that overcomes this limitation. We first observe that nonlinear operations in the decode stage exhibit heterogeneous workload patterns, which can be handled effectively via a hybrid approach. We then realize this with two key contributions: 1) an adaptive segmented lookup-table protocol based on TFHE that enables efficient and accurate evaluation of nonlinear operations; and 2) a scheme-aware operator-selection framework that automatically assigns each nonlinear operator to CKKS or TFHE to minimize end-to-end decoding latency. We demonstrate that ROSETTA achieves up to $4.8\times$ Softmax speedup and $1.5$--$2.1\times$ end-to-end speedup over the SOTA framework CacheMir.
Odin is the first open-source end-to-end GPU CKKS implementation of Llama-3, an FHE inference system that co-designs ciphertext packing and model execution for Llama and uses a feature-major cross-layer layout to unify residual connections and layer interfaces.
Yu-Hang Fan, Yu-Si Chen, Kan-Yu Ye et al.· 0 citations
This work presents a distributed inference framework that integrates speculative decoding across edge and cloud, and shifts the bulk of computation to the edge, significantly lowers inference time and cloud cost, and preserves the accuracy of the big model without any retraining requirement.
D. J. Bajpai, K. Upadhyay, M. Hanawal· 0 citations
Cloud-based Large Language Model (LLM) inference services typically require users to submit plain-text inputs, thereby posing severe privacy risks. Existing privacy-preserving paradigms are mostly task-specific and often necessitate pervasive modifications to the entire server-side model. This reliance introduces subst...
Wentao Zhong, Yu-Ting Li, Di-Cong Yu et al.· IEEE Transactions on Informa...· 0 citations
CoVeil is proposed, a defense mechanism which dynamically optimizes transmitted signals to suppress leakage during decoding time while preserving the collaborative quality, and consistently improves the privacy-utility trade-off over existing baselines by reducing data leakage.
Ke-Jia Zhang, Tianyuan Zou, Zi-Xuan Gu et al.· 0 citations
With the increasing application of machine learning in sensitive domains such as healthcare, the demand for privacy-preserving techniques in decision tree inference has grown urgent. Traditional decision trees, when executed in untrusted environments, pose a risk of exposing sensitive data to untrusted parties. Therefo...
Shi-Wen Wei, Zhi-Li Chen, Yan-Zhuo Yu et al.· IACR Transactions on Cryptog...· 0 citations
Fully Homomorphic Encryption (FHE) enables computation on encrypted data, preserving privacy throughout analysis. While its privacy is very strong, FHE is much slower to execute than the original computation. In particular, due to the recent success in accelerating its compute, the performance bottleneck shifts to the...
A. W. B. Yudha, Erwin Eko Wahyudi, R. Rajagede et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.