"Physics-Informed Distributed Model Predictive Control for Thermally Aware Energy Management in AI Data centers"
Abstract
"This dataset and reproducibility package accompanies the research paper \"Physics-Informed Distributed Model Predictive Control for Thermally Aware Energy Management in Hyperscale Artificial Intelligence Data Centers,\" provisionally accepted as a Regular Paper in the IEEE Transactions on Industrial Informatics (Paper No. TII-26-5675.R1).Hyperscale Artificial Intelligence (AI) data centers present severe spatial-temporal power and thermal volatility driven by bursty multi-megawatt tensor workloads, tight server thermal limits, and volatile wholesale electricity markets. This package contains the complete numerical simulation suite, trained neural-physics surrogate models, curated real-world telemetry, and physical hardware-in-the-loop (HIL) execution logs for a two-tier, distributed, physics-informed model predictive control (PI-DMPC) architecture.The package includes:1. Curated Exogenous Telemetry: 336 hours (two 168-hour representative winter and summer stress windows: January 1\u20137, 2025 and August 1\u20137, 2025) of calibrated 30 MW hyperscale facility telemetry, encompassing real Alberta Electric System Operator (AESO) wholesale electricity pool prices, 15-minute grid marginal carbon emission intensity, dry-bulb\/wet-bulb ambient weather data from the Calgary Eagle station, and empirical compute-cluster trace workloads.2. Physics-Informed Neural Network (PINN) Model: Pre-trained PyTorch surrogate weights (neural_physics_solver.pt) enforcing thermodynamic energy balance constraints for rack- and room-level temperature evolution under non-uniform IT workloads.3. Closed-Loop Simulation Suite: Python\/CVXPY\/HiGHS benchmark drivers implementing PI-DMPC via the Alternating Direction Method of Multipliers (ADMM), alongside 9 comparative baselines (B1\u2013B9) including classical heuristic control, centralized linear\/nonlinear MPC, decentralized MPC, continuous Twin Delayed DDPG (TD3) deep reinforcement learning, and an acausal clairvoyant Oracle-MPC.4. Physical OPAL-RT OP5600 HIL Logs: Deterministic real-time companion controller execution logs, Simulink\/RT-LAB models (PI_DMPC_OP5600_console_HIL.slx), and 25-channel sensor\/actuator traces recorded at a 10 ms hard-synchronized target step across 40,320 target iterations (336 hours of operation), achieving numerical equivalence with simulation at a maximum absolute error of 2.22e-10.All reported figures, tables, and constraint satisfaction metrics (100% service latency satisfaction and zero thermal violations across 100-sample Monte Carlo runs) can be verified directly using the included automated audit scripts. "