Jul 2026
Layer-Parallel Inference Reduces Encrypted Nonlinear Depth in Transformers
Ablations show that softmax approximation dominates the error budget and CKKS arithmetic noise is negligible in the authors' setting, suggesting that SNLP is complementary to block-level FHE-friendly operator design rather than a replacement for it.
Ligong Han, Kai Xu, Hao Wang et al.
· arXiv.org · 0 citations