Rethinking Attention: Entropy-Guided Iterative Refinement for Language Models
Standard Multi-Head Attention (MHA) computes contextual representations in a single forward pass, providing no mechanism to detect or correct diffuse or suboptimal attention distributions. Existing efforts to improve attention have primarily targeted computational efficiency or sequence length, leaving attention qualit...