Every recurrence of a looped language model adds cost in training, decoding, prefill, and reinforcement learning (RL). The closer recurrent states get to fixed points, the less the path to them matters. This enables truncated backpropagation in training; terminal key-value (KV) sharing for decoding with almost no loss...
Benhao Huang, Chu-Fan Shi, Jun-Lin Chen et al.· 0 citations
BF16 is now standard in large-scale pretraining, including in fused attention kernels such as FlashAttention, and these kernels are widely trusted. When we used FlashAttention-3 to pretrain a 450M-parameter transformer on 50B tokens, however, we ran into a problem: training was healthy for 25B tokens, then the gradient...
Jun-Lin Chen, Daize Dong, Huan-Wei Di et al.· 0 citations
This work designs Diamond Agent, an agentic system that enables intelligent execution of HPC workflows across heterogeneous clusters with typed skills as the interface, and reduces the median additional completion time relative to the fastest observed placement from 42 seconds to 4 seconds, a 10.5x reduction.
Haotian Xie, Jun-Lin Chen, Ming-Kai Zheng et al.· 0 citations
Cremes is proposed, an adaptive and cost-efficient scaling framework that ensures microservice recovery within the spot instance grace period and maintains SLO violation rates under preemptible environments below 6.7%.
Liao Chen, Chenyu Lin, Junlin Chen et al.· IEEE International Symposium...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.