A safety-constrained deep reinforcement learning framework for source–load–storage coordinated operation of a grid-connected green data center and shows how separating reward learning, cumulative safety pricing, and one-step engineering projection changes low-carbon dispatch within the specified model.
Abstract
Green low-carbon data centers operate as coupled cyber-energy systems whose dispatch must coordinate renewable generation, grid exchange, battery storage, cooling load, flexible computing workload, carbon-intensity signals, and reliability constraints. This study develops and evaluates a safety-constrained deep reinforcement learning framework for source–load–storage coordinated operation of a grid-connected green data center. The operating problem is formulated as a constrained Markov decision process with state variables describing the IT load, deferrable workload backlog, renewable availability, electricity price, marginal carbon intensity, battery state of charge, server-room temperature, reserve margin, and calendar context. The action space covers grid import and export, renewable utilization, storage charge and discharge, workload shifting, and cooling control. The learning architecture combines a constrained actor–critic policy, adaptive Lagrangian safety critics, and a control barrier function (CBF)-based action shield that projects unsafe actions onto an explicitly defined operating set before plant execution. The shield is specified as a low-dimensional quadratic projection over state-dependent SOC, thermal, reserve, SLA, and grid-interface constraints, while cumulative risks are priced through Lagrangian safety budgets during policy training. The evaluation uses a controlled and auditable benchmark simulation with normalized public-data-compatible profiles, declared scenarios, random seeds, neural-network settings, and mechanism-matched baselines; it is not a telemetry-based verification or hardware certification of a deployed data center. Within this declared benchmark, the proposed safe DRL controller produces a simulated 13.1% emission reduction relative to the Rule-based controller, 95.8% renewable utilization, a normalized annual cost of 0.91, and fewer boundary contacts than the tested unconstrained, Lagrangian-only, and shield-only PPO variants. These percentages are simulator outputs relative to the stated benchmark and must not be interpreted as measured field savings. The results show how separating reward learning, cumulative safety pricing, and one-step engineering projection changes low-carbon dispatch within the specified model.
A deep reinforcement learning (DRL) framework that jointly co-schedules computing and thermal resources so that a hyperscale data center can operate as a grid-interactive flexible load and supports the evolution of hyperscale data centers from passive electricity consumers toward active, grid-interactive participants in renewable-penetrated power systems.
High renewable penetration and large-scale green hydrogen production are accelerating the formation of the new-type power system (NTPS), in which electrical dispatch, electrolysis, hydrogen storage, fuel-cell reconversion, and flexible demand must be coordinated under nonlinear network physics and uncertain renewable, load, and hydrogen-demand trajectories. This study develops a physics-informed distributionally robust multi-agent reinforcement learning (PI-DRO-MARL) framework for coordinated NTPS operation with integrated electricity–hydrogen coupling. The operational objective is to minimize worst-case expected operating cost, including generation and grid-exchange cost, electrolysis and hydrogen-delivery cost, storage degradation, renewable curtailment, and load- or hydrogen-shedding penalties, while satisfying AC power-flow balance, voltage limits, line-loading limits, ramping limits, battery state-of-charge constraints, hydrogen-storage dynamics, and electrolysis/fuel-cell conversion constraints. The framework embeds physics-informed residuals and projection operators into a centralized-training decentralized-execution architecture; represents renewable, electrical-load, hydrogen-demand, and price uncertainty through statistically calibrated Wasserstein ambiguity sets; and trains agents with robust value estimation and feasibility-aware action correction. Validation is conducted on a modified IEEE 33-bus distribution network coupled with a 12-node hydrogen system, with additional scalability checks on modified IEEE 69-bus and IEEE 123-node reference systems. Across ten random seeds, the primary case shows an operating cost of USD 8850 with a 95% confidence interval of USD 8770–8940, a mean constraint-violation rate of 0.37%, and a shifted-scenario cost increase of 12.6%, outperforming deterministic optimization, stochastic programming, standard reinforcement learning (RL), proximal policy optimization (PPO), soft actor–critic (SAC), multi-agent deep deterministic policy gradient (MADDPG), constrained RL, safe RL, and robust RL baselines. Ablation, Wasserstein-radius, time-step, and stress-test analyses further show that distributional robustness, physics-informed projection, and multi-agent coordination provide distinct and complementary benefits. The results support PI-DRO-MARL as a simulation-validated architecture for real-time, uncertainty-aware NTPS dispatch, while field deployment still requires digital-twin calibration, hardware-in-the-loop testing, and site-specific operational validation.
Fei Liu, Outing Zhang, Jun Yin et al.· Energies· 0 citations
This paper proposes an edge-cloud collaborative physics-informed reinforcement learning framework for production data center HVAC control that integrates a physics-informed cold-start solution using Adaptive Particle Swarm Optimization, a three-time-scale edge–cloud architecture, and a constraint-aware safe projection layer that embeds thermal safety hard constraints directly into the neural network policy.
Shichao Huang, Yi-Bing Zhou, Yuan Liu· Italian National Conference...· 0 citations
This paper formulate networked grid operation as a constrained decentralized partially observable Markov decision process and proposes a safe multi-agent collaborative learning framework that aims to reduce operating cost, load shedding, renewable curtailment, and carbon-relevant corrective burden.
Jia-Yi Zhang, Bing Fang, Huan-Xiu Xiao et al.· International journal of pat...· 0 citations
The increasing penetration of weather-driven renewable energy sources in smart grids introduces operational instability, harmonic distortion, and elevated switching costs due to the limitations of rule-based and deterministic control strategies. This study proposes a deep reinforcement learning-based adaptive switching framework to enhance renewable utilization while minimizing operational risk and economic cost. A simulation-derived dataset incorporating renewable generation, load demand, total harmonic distortion, voltage deviation, frequency variation, and risk–cost indices was generated from a risk–cost optimized smart grid model and implemented in Google Colab. The switching problem was formulated as a Markov decision process with a state space composed of power quality and economic variables, and a discrete action space representing operational modes. A Deep Q-Network agent was trained over 24-hour episodes to learn optimal switching policies. Comparative evaluation against conventional, rule-based, and analytical risk–cost optimization strategies demonstrated up to 14% reduction in total harmonic distortion, 22% reduction in switching frequency, 17% reduction in operational cost, and 19% improvement in risk mitigation, while increasing renewable penetration by 11%. The proposed framework provides a scalable and intelligent solution for industrial smart grid applications.
M. Meyyappan, P. Avirajamanjula, P. Marimuthu et al.· ITEGAM- Journal of Engineeri...· 0 citations
In grid-connected smart integrated energy systems with high shares of renewable generation, source-side variability and inadequate coordination among battery storage, other energy carriers, and the external grid limit local renewable-electricity utilization and impede deep decarbonization. This study proposes a machine-learning-assisted, renewable-driven framework for multi-energy coupling and scenario-based multi-objective optimization of electricity–heat–hydrogen–storage systems. Historical meteorological and load data are processed using K-means clustering and Latin hypercube sampling to construct representative operating scenarios across multiple volatility regimes and characterize source–load uncertainty. The equipment model includes photovoltaic arrays, wind turbines, heat pumps, electrolyzers, fuel cells, grid-interactive battery energy storage, thermal storage, and hydrogen storage; cross-carrier conversion dynamics and emissions from purchased electricity and natural gas are embedded in the energy-balance constraints. A mixed-integer linear programming formulation then co-optimizes battery charging and discharging, grid exchange, and other multi-energy flows with respect to operating cost, carbon emissions, and renewable-energy curtailment. At 95% renewable-energy penetration, the proposed method achieves a renewable-energy absorption rate of 91.6% and a curtailment rate of 8.4%. Across the carbon-price cases, annualized operating cost ranges from 126.5 × 104 to 141.2 × 104 USD yr−1, while carbon-emission intensity ranges from 26.4 to 38.5 gCO2/kWheq. Under the specified high-risk grid disturbances, the coordinated strategy limits load shedding to 1.8%—73% below deterministic scheduling and 79% below the heuristic benchmark—and maintains 92.6% hydrogen self-sufficiency. These results provide a data-driven modeling and decision framework for battery–grid coordination and deep decarbonization in smart integrated energy systems.
Yao Tong, Hai-Ling Ma, Fu-Yi Du· Batteries· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.