SAGE is proposed, a noise-aware shrinkage method that adaptively attenuates privatized estimates according to their estimated signal quality, and shows that shrinkage reduces the quadratic update-risk term faster than the linear descent term, preserving useful descent while limiting the influence of noise-dominated updates.
Abstract
Differentially private zeroth-order optimization (DP-ZO) enables memory-efficient private fine-tuning of large language models using only forward evaluations. Existing aggregation-based DP-ZO methods reconstruct model updates at a fixed scale, ignoring that the strength of useful signals varies throughout training. Consequently, noise-dominated updates may receive excessive weight and degrade model utility. To address this issue, we propose SAGE, a noise-aware shrinkage method that adaptively attenuates privatized estimates according to their estimated signal quality. SAGE subtracts the known Gaussian noise variance from the observed second moment to estimate the underlying signal energy, stabilizes this estimate through temporal tracking, and compares its current signal-to-noise level with a warm-up reference to derive a bounded shrinkage factor. As pure post-processing, SAGE requires neither additional privacy budget nor model queries and introduces only constant additional state. Our theoretical analysis shows that shrinkage reduces the quadratic update-risk term faster than the linear descent term, preserving useful descent while limiting the influence of noise-dominated updates. Experiments on RoBERTa-large, OPT-1.3B, and OPT-6.7B demonstrate that SAGE outperforms existing baselines in most settings under the same privacy budgets while preserving the forward-only memory efficiency of DP-ZO.
This work introduces RAPTOR - a Role-Aware Private Training framework, which alternates shared and expert optimization and targets each failure directly, using expert-specific clipping and noise together with a public expected-owner denominator and a count-independent update schedule that avoids conditioning on private...
Duc Dm, Khai Le-Duc, D. Nguyen et al.· 0 citations
Zeroth-order (ZO) optimization with SGD in random subspaces enables memory-efficient fine-tuning of large language models without backpropagation. However, high gradient estimation noise fundamentally undermines adaptive optimizers like Adam. We propose SubZero+, which achieves practical adaptive ZO optimization throug...
Zi-Ming Yu, Shu-Yao Xiao, Xingyu Zhao et al.· 0 citations
Under homoscedastic retrieval noise, it is shown that retrieval failure decays exponentially with task separation relative to noise, and explicit finite-sample conditions under which RAG-FT achieves lower risk than both target-only and full-corpus training are derived.
Debias-SparseGPT is introduced, a post-training pruning method incorporating representational debiasing using a second-order term defined over demographically contrasting inputs that consistently reduces pruning-induced bias compared to SparseGPT while preserving model perplexity and zero-shot accuracy.
Irina Proskurina, Guillaume Metzler, Antoine Gourru et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.