Fiona: Accelerating FHE Inference with Packing-Aware Ternary Weights
FIONA is presented, an offline optimizer that selectively ternarizes weights within a given packing layout based on the estimated effect of ternary conversion on the model's performance, and fits lower-degree replacements under a cumulative accuracy budget, reducing multiplicative depth and bootstrapping.