DeepShare is a scheduler that uses a continuous tenant-assurance signal to coordinate these decisions at runtime to achieve a more advantageous utilization-QoS trade-off than optimizing quotas, scheduling, and resource sharing independently.
Jing-Hao Wang, Yi-Hang Zhou, Xiao Zhou et al.· 0 citations
Concurrent multi-agent workflows expose future dependencies and serving-state requirements while running on heterogeneous GPU pools with time-varying load, model residency, and resource availability. The logical workflow defines the required computation, whereas its physical scheduling units, model-lifecycle actions, r...
Jing-Hao Wang, Yi-Feng Zhang, Xiao Zhou et al.· 0 citations
This work presents SpecBox, a runtime built around speculative sandbox preallocation tailored for dynamic LLM agent execution pipelines, and implements keyword matching and streaming semantic embedding to enable intent-driven sandbox prewarming, which identifies pending tool execution demands mid-LLM token generation a...
Yi-Hui Zhang, Tian-Yu Wo, Jing-Hao Wang et al.· arXiv.org· 1 citation
ElastiCo is presented, an elastic co-location framework that enables training and inference workloads to safely share GPUs through three integrated mechanisms that decomposes the resulting multi-resource allocation problem into per-job configuration selection subproblems via dynamic per-resource shadow prices.
Jing-Hao Wang, Yi-Hang Zhou, Xiaoyang Sun et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.