Sliding-window beats linear attention
This work shows that Sliding Window Attention (SWA) with sinks performs as well or better than post-trained Linear Attention models, and recommends switching to SWA instead of post-training linear models.
Alexia Jolicoeur-Martineau, R. Sukthanker, Pashmina Cameron et al.
· 0 citations