Rotary Position Embedding (RoPE) encodes token positions by rotating each two-dimensional channel of the query and key vectors at a channel-specific frequency, making the attention logits invariant to a common shift of positions. However, this rotation is periodic, and it leads to position aliasing where relative posit...
Guan-Cheng Du, Luo-Tian Huang, Shao-Wen Wang et al.· 0 citations
Language model pretraining has become almost synonymous with prohibitive cost, placing it out of reach for much of the academic and open-source communities. Although strong open-source efforts already exist, including open-weight models and open-source training recipes, a cost-efficient, hardware-accessible, and open-s...
Kai-Rong Luo, Jia-Rui Cui, Yao-Rui Yin et al.· 0 citations
This report presents an open pretraining recipe that trains a collection of Puro-2B models from scratch on up to 1.4 trillion tokens with FP8 precision on consumer-grade RTX 5090 GPUs, and derives a Puro Cost Scaling Law that relates training cost to average model performance.
Kairong Luo, Jia-Rui Cui, Yao-Rui Yin et al.· 0 citations
2D-RoPE is introduced, which organizes text into a 2D grid rather than a 1D sequence and assigns each token a row ID and a column ID, and suggests that viewing text in 2D can benefit language modeling.
Haodong Wen, Yiran Zhang, Yingfa Chen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.