Multi-granularity collaborative scheduling method for distributed AI training in all-optical data center networks
Abstract
The rapid growth of large-scale AI models has driven the emergence of AI data centers (AIDCs), where distributed training produces massive communication demands under multi-job concurrent execution. All-optical data center networks have emerged as a promising solution due to their high bandwidth and low latency. However, the diverse communication demands introduce severe network resource contention. To address this issue, we propose a multi-granularity collaborative scheduling (MGCS) method for distributed training workloads over an all-optical switching network architecture. It integrates optical time-slot switching (OTS) and optical circuit switching (OCS), enabling flexible allocation of network resources to accommodate diverse communication demands. MGCS follows a phase-aware layered scheduling logic. For pipeline-parallel (PP) communication, it first formulates the deterministic OTS scheduling problem as a combinatorial optimization problem and develops both a mixed-integer linear programming (MILP) model and a heuristic algorithm. It then schedules data-parallel (DP) communication through a hybrid multi-granularity scheduling method that dynamically coordinates OTS and OCS resources. We build a small-scale all-optical switching network testbed and conduct large-scale simulations to evaluate the proposed method. The results indicate that MGCS achieves up to a 30.60% reduction in epoch training time and a 49.70% reduction in the network resource occupation ratio.