Jacobian-guided Noise Injection for Quantization Robustness in Large Language Models
This work proposes Jacobian-Guided Noise Injection, a training strategy that injects zero-mean Gaussian noise into pre-attention logits, with variance derived directly from the Jacobian Frobenius norm, which provides a way to identify the optimal noise variance based on the local attention sensitivity.