Jul 2026
Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks
Mask2Shield (M2S), a masked-forward alignment method that trains a model under this functional pruning procedure, reduces successful recomputed pruning attacks from 80--279 to 1--44 out of 313 prompts while generally preserving four capability benchmarks.
Jincheng Ying, Ming-Hui Xu, Yinhao Xiao et al.
· arXiv.org · 0 citations