A key bottleneck in 3D Gaussian Splatting training is the continual growth of Gaussian primitives, which increases optimization cost and slows convergence, especially at high resolutions. We propose Laplacian Frequency Hierarchies, a simple yet efficient 3DGS scheme that combines Laplacian image decomposition with coar...
Yixiong Yang, Sirius Z. Zhang, Q. Yan et al.· 0 citations
Communication poses a dominant bottleneck in distributed data parallel training with synchronous stochastic gradient descent, whereas the traditional AllReduce collective used for gradient synchronization limits the efficient utilization of communication compression strategies. In this paper, we propose a low-bit and s...
Jia-Qi Li, Shao-Huai Shi, Jing Peng et al.· Proceedings of the Internati...· 0 citations
LeanGRPO is presented by restructuring the data-parallel layout and introducing two recompute-free training schedules for trajectory-logprob diffusion RL, which achieves up to 1.83x end-to-end speedup while preserving the original optimization objective.
Si-Jie Wang, Zhi-Qiang Tan, Xin-Rui Yang et al.· 0 citations
BiDiRL, a hybrid time-space multiplexing architecture for asynchronous, disaggregated RL designed to reduce resource idleness, is presented, including a hot-switch runtime that enables rapid switching between rollout and training resources with negligible overhead and a static, scheduling-aware planner based on time-pe...
Zhiqiang Tan, Maoxin Wang, Sijie Wang et al.· 1 citation
Xema is presented, a memory-efficient diffusion serving system that exploits predictable tensor lifetimes for trace-guided memory optimization and introduces an offline planner that jointly selects parallelism, concurrency, and memory control under GPU memory and SLO constraints.
Xueze Kang, Guangyu Xiang, Suyi Li et al.· arXiv.org· 1 citation
KernelFlume is presented, a decode-centric architecture that disaggregates the stable projection/FFN path from core-attention computation: weight nodes execute dense projection/FFN kernels, while weightless attention nodes store token-range KV partitions and scale with request-state demand.
Guangyu Xiang, Xueze Kang, Lin Zhang et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.