Preprint
Aug 2026
Beyond Capacity: Scalable MoE LLM Inference via High-Bandwidth Flash with Direct GPU and HBM Paths
This work explores an HBF organization that simultaneously exploits two independent expert-delivery routes: a direct path that transfers expert weights from HBF to the GPU and a relay path that transfers them from HBF through the HBM base die to the GPU.
Seeyeon Kim, Juhyeong Jin, Joo-Young Kim
· 1 citation