Jul 2026
Prox: Training-Free FFN Activation Sparsity via Approximate Intermediate-Channel Salience in LLMs
Prox is a two-stage training-free framework for sparse SwiGLU FFNs that outperforms training-free baselines at all sparsity levels, achieves up to a $1.99\times end-to-end decoding speedup at 70\% FFN sparsity, and is compatible with quantization and sparse attention.
Jinyi Liu, Wei Chen, Pengyu Chen et al.
· arXiv.org · 0 citations