What Softmax Throws Away: Mass-Aware Attention for Evidence Accumulation
Mass-Aware Attention is proposed, which generalizes standard L1 normalization to an Lp family and is positioned as a general normalization principle for improving predictor-facing representation informativeness by controlling repetition invariance in standard attention.