Skip to content

Author

Yichun Yin

We have 2 of 29 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Switch Attention: Towards Dynamic and Fine-grained Hybrid Transformers

SwiAttn is a novel hybrid transformer that enables dynamic and fine-grained routing between full attention and sliding window attention, and dynamically routes the computation to either a full-attention branch for global information aggregation or a sliding-window branch for efficient local pattern matching.

Yusheng Zhao, Hourun Li, Bohan Wu et al. · 4 citations
Preprint Aug 2026

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction

AdaMTP is proposed, an adaptive training paradigm that dynamically aligns the prediction horizon with the intrinsic predictability of the sequence, and consistently outperforms standard MTP in both task performance and inference speedup.

Ziqiang Cui, Han Shi, Bowei He et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.