A heterogeneous-graph Proximal Policy Optimization (PPO) scheduler in which the placement head is permutation-equivariant, the delay head is permutation-invariant, and the parameterization stays invariant to the processor count N.
Abstract
As data-center interconnects move to co-packaged optics (CPO), high-power application-specific integrated circuits (ASICs) and heat-sensitive optical engines share a single interposer, and the resulting intra-module thermal coupling overwhelms conventional schedulers. Thermal-aware microservice directed acyclic graph (DAG) scheduling on CPO modules is recast here as a question of symmetry. Whereas a homogeneous-graph policy assumes the full node-permutation symmetry, that symmetry is broken twice: by the distinct task and processor node types, and by the asymmetric ASIC–engine coupling. We therefore propose a heterogeneous-graph Proximal Policy Optimization (PPO) scheduler in which the placement head is permutation-equivariant, the delay head is permutation-invariant, and the parameterization stays invariant to the processor count N. Because these symmetries hold by construction, the policy transfers zero-shot across module sizes. Heterogeneous edge typing and the resistor–capacitor (RC) coupling edge attribute are isolated by a six-test ablation chain. Evaluated on the Alibaba 2021 microservice trace across all module sizes and ambient regimes under the standard auto-cool budget, the proposed scheduler cuts the thermal-violation rate from roughly 98% under the Heterogeneous Earliest-Finish-Time (HEFT) heuristic to about 0.3%; at the hot operating point it lowers peak temperature by about 25% and raises DAG completion from about 26% to 100%, with the rare residual violations most frequent in the extreme-ambient band. With the env auto-cool budget disabled, a controlled single-axis comparison shows that removing the RC-coupling edge attribute raises the violation rate by over an order of magnitude, isolating its contribution. A single parameter set serves every N without retraining.
Compute-in-memory (CIM) architectures mitigate the von-Neumann bottleneck by embedding computation directly within memory crossbars, delivering order-of-magnitude improvements in energy efficiency and throughput. Among the emerging technologies, non-volatile-memory (NVM)-based CIM is particularly attractive owing to it...
Can Gao, Xuejin Li, Kaiwei Zou et al.· IEEE Non-Volatile Memory Sys...· 0 citations
Post-placement optimization has been actively studied using techniques such as gate sizing, threshold-voltage (Vth) swapping, and buffer insertion. Recently, many approaches have adopted machine learning, particularly reinforcement learning (RL), to automate these optimizations. However, existing RL-based methods fall...
Kijung Kong, Heechun Park· International Symposium on L...· 0 citations
PReCCL is a drop-in NCCL replacement that combines software inband telemetry with cross-VT workload reallocation, and implements in-band monitoring within the CCL, and precisely measures the stall counts of each VT, and piggybacks the telemetry meta-data on existing collective traffic.
Zhiyong Chen, Kaihui Gao, Li Chen et al.· Conference on Applications,...· 0 citations
ReTri, a bidirectional All-to-All schedule for ORNs based on the Trivance algorithm is presented, a bidirectional All-to-All schedule for ORNs based on the Trivance algorithm that improves completion time and improves reconfigurable Bruck by up to 2.1×.
Anton Juerss, Stefan Schmid· Conference on Applications,...· 0 citations
Pack sizing for competition and road hybrids is a constrained energy-balance problem whose feasible set shrinks once communication integrity, command latency, and solar assist are admitted as coupled limits. Cell-material and management reviews show that fade, thermal margin, and leftover capacity after a duty cycle be...
G. Aa, S. S· International Journal of Cre...· 0 citations
Optical-circuit-switched interconnects have become one of the core components for AI training due to their flexible topology reconfiguration. In contrast to the applications carried by traditional data center networks, large-scale language model training is highly sensitive to network failures, where frequent disruptio...
Liang Qin, Xing-Yu Liu, Wen-Ting Wei et al.· IEEE Transactions on Cogniti...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.