A serving system that treats the diffusion block as a compilation unit that improves end-to-end execution time by up to 2.7 times over the strongest surviving baseline under the same 8-GPU placement and remains feasible at the largest batch sizes where multiple baselines run out of memory, while preserving task quality relative to the dense reference.
Jia-Nian Zhu, Hang Wu, Ying-Hui Li et al.· 0 citations
Prism abstracts the highly dynamic diffusion workload into a predictable, static execution flow transparent to the compiler, and achieves this via three techniques: spatial regularization, temporal stabilization, and specialized kernels that selectively bypass padding data.
Jianian Zhu, Hang Wu, Yinghui Li et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.