Computing-And-Power Cooperative Scheduling for Distributed AI Training at Scale
As AI training moves toward geo-distributed elastic infrastructures, its bottlenecks extend beyond GPU availability to variable energy and network conditions. The key challenge is translating heterogeneous placement, scaling, congestion, and training-stage effects into useful progress. We propose Effective Compute Util...