Skip to content
Conference

Computing-And-Power Cooperative Scheduling for Distributed AI Training at Scale

Jul 2026 · Fall Joint Computer Conference · pp. 57-64 · 0 citations · 21 references

Abstract

As AI training moves toward geo-distributed elastic infrastructures, its bottlenecks extend beyond GPU availability to variable energy and network conditions. The key challenge is translating heterogeneous placement, scaling, congestion, and training-stage effects into useful progress. We propose Effective Compute Utility (ECU), an interpretable progress signal, and build ECU-Aware, an online scheduler that coordinates data-center placement, elastic GPU allocation, and execution timing under budget and deadline constraints. Using Lyapunov optimization, ECU-Aware balances operating cost against long-term progress. Experiments with real traces and a small-scale prototype show that ECU-Aware improves the cost-latency-SLA tradeoff. In a 500-job workload, it reduces cost by 48.9 percent relative to a static policy, lowers average JCT by 36.0 percent, and decreases deadline misses by 17.2 percentage points compared with practical baselines.

View source

Similar papers

Conference Aug 2026

AI-Driven Autonomous Resource Scheduling in Cloud–Edge Environments

Cloud Edge computing provides services with low latencies by allocating workloads between centralized cloud servers and geographically close edge nodes, but dynamic workloads and nonhomogeneous resources render the scheduling of workloads a thorny multi-objective optimization problem. The given paper proposes a framewo...

Vaibhav Sawalkar, Anuja Gaikwad, Vishal Bogum et al. · 0 citations
Open access Sep 2026

Sustainable AI Request Scheduling with Joint Compute, Network, and Power Optimization

RAPID is proposed, a region-aware and power-informed scheduling framework that integrates static and online heuristic schedulers for large-scale AI request scheduling that significantly reduces carbon emissions, electricity costs, and total energy consumption.

B. Ding, Cai-Ning Wang, Ka-Fei Tang et al. · 0 citations
Conference Aug 2026

Predictive Versus Projection-Based Scheduling for O-RAN Fronthaul Deadline Assurance

In the Open Radio Access Network (O-RAN) 7.2x split, the Distributed Unit (O-DU) must finish upper physicallayer processing before strict fronthaul deadlines. However, heterogeneous acceleration through CPU fallback, shared acceleration, and dedicated acceleration introduces profile-selection, contention, and queueing...

Viet Anh-Huy Le, Anh-Tien Tran, Thanh Thien-An Dang et al. · 0 citations
Preprint Sep 2026

DLB: Distributed Load Balancing at Scale for Generative AI Inference

DLB, the Distributed Load Balancer is introduced, a novel system designed to minimize end-to-end user latency for large-scale, heterogeneous workloads and its design choices and practical experiences gained from the system in production are detailed.

S. Balseiro, B. Wydrowski, Sameer Agarwal et al. · 0 citations
Open access 2026

JET: Decoupled Placement and Online Coevolutionary Scheduling in Data-Intensive Mobile Edge Computing

Data-intensive Mobile Edge Computing (MEC) demands a synergy between server placement and task scheduling, yet these processes operate on disparate optimization horizons. This paper proposes JET, a two-stage optimization framework that bridges this gap through placement-aware scheduling. In the offline stage, an evolut...

Sadra Galavani, Abolfazl Younesi, Mohsen Ansari · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.