Reliable FP8 attention remains a barrier to fully native 8-bit large language model training. We derive how forward-backward inconsistencies produce stale delta and empirically show how it distorts training dynamics. Our stale-delta hybrid runs show a modest loss gap at 569M parameters but substantial loss increases an...
Hao-Zhan Tang, Hao Kang, Han Cai et al.· 0 citations
LeapQuant is proposed, a training-free method that achieves near-lossless performance under 8-bit recurrent-state quantization and substantially reduces memory and compute costs during inference.
Frontier open-weight models are increasingly available, but serving them still largely assumes datacenter infrastructure. We present FreeToken, an edge-native MoE serving system that treats a personal machine not as a small GPU, but as a unified, elastic inference platform. FreeToken co-designs the full serving stack,...
Shuo Yang, Xiao-yun Fan, Melissa Z. Pan et al.· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.