DeepShare is a scheduler that uses a continuous tenant-assurance signal to coordinate these decisions at runtime to achieve a more advantageous utilization-QoS trade-off than optimizing quotas, scheduling, and resource sharing independently.
Abstract
Multi-tenant GPU clusters frequently remain underutilized even when tenants experience long queueing delays, because quota control, queue ordering, preemption, and GPU sharing are driven by different local signals. We present DeepShare, a scheduler that uses a continuous tenant-assurance signal to coordinate these decisions at runtime. DeepShare combines elastic quota borrowing, tenant-specific runtime prediction, cost-aware best-effort preemption, and interference-aware MPS colocation, while using the same assurance signal to decide when borrowed capacity should be reclaimed and when sharing should become more conservative. In trace-driven experiments on 23,859 Venus jobs and 3,200 internal jobs, DeepShare achieves an average GPU utilization of 70.58%, a 29.5% improvement over the strongest non-intrusive sharing baseline, while reducing average queueing delay by 46%. On a 16-GPU Kubernetes testbed, it reduces the average job completion time by 34% and maintains 93% QoS compliance for guaranteed tenants. These results show that treating tenant assurance as a runtime control loop achieves a more advantageous utilization-QoS trade-off than optimizing quotas, scheduling, and resource sharing independently.
SIF is introduced, a metric built on the notion of partial-nodes that is independent of historical workload knowledge that matches both theory and production and COMPASS-ABS is proposed, which employs the COMPact-ASSured (COMPASS) algorithm to confine the cluster state within a tight Anchor-Based Space (ABS).
Sharing GPUs among many deep learning models is crucial for cost-efficient inference, but bursty multi-model workloads can easily overwhelm GPU capacity, causing severe tail latency and SLO goodput drops. Existing solutions—whether traffic-aware scheduling or hardware-level resource partitioning—can only juggle content...
Shi-Jie Peng, Yanying Lin, Cheng-Zhi Lu et al.· Proceedings of the Internati...· 0 citations
LLMVisor is presented, a roofline-guided latency attribution model that captures the memory-bound and compute-bound phases via a concise piecewise-linear form over features proportional to FLOPs and memory I/O traffic and runs efficiently at microsecond scale.
Shuowei Jin, Xue-Shen Liu, Jiaxin Shan et al.· 3 citations
Large Language Models (LLMs) are increasingly deployed in latency-sensitive applications, where real-time serving must satisfy stringent service-level objectives (SLOs). However, request intensities fluctuate over time, and under low load LLM services leave a substantial fraction of GPU task idle. A promising approach...
SemaCredit is presented, a receiver controller that admits each remote-memory operation against a vector of target-resource demands and returns each component when its corresponding HBM, Atomic, or response stage completes, reducing small-operation P99 latency by 52.4% under Atomic contention and 10.2% under response i...
Fan Yang, Jiaqi Liu, Tao Jiang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.