Jul 2026
Beyond Prefill-Decode Disaggregation: Dissecting LLM Inference for Heterogeneous Platforms via Dynamic Operator Scheduling
DOPS (dynamic operator scheduling), a hardware-aware, closed-loop framework that jointly optimizes operator scheduling and blockwise weight layouts and supports systematic analysis of workload sensitivity and hardware scalability for LLM serving is presented.
Jiaqi Yang, Jia-Yi Li, Yihan Fu et al.
· arXiv.org · 0 citations