Jul 2026· Advances in Data Science and Adaptive Analysis· Vol 18· 0 citations
TL;DR
Convergence and multi-run statistical analysis further confirm the robustness, stability, and reproducibility of the trained policy, demonstrating the effectiveness of DRL as an intelligent and scalable solution for next-generation microgrid PQ control.
Abstract
Hybrid microgrids integrating photovoltaic (PV) arrays, wind tur-binedriven PMSG units, fuel cells, and battery storage enhance sustainability but face serious power quality (PQ) challenges due to the intermittent and nonlinear behavior of renewable sources and loads. Traditional PI, PR, and hybrid intelligent controllers offer acceptable nominal performance but lack adaptability and predictive capability under rapidly varying disturbances. To overcome these limitations, this paper proposes a Deep Reinforcement Learning (DRL) frame-work based on the Twin-Delayed Deep Deterministic Policy Gradient (TD3) algorithm for real-time PQ management, where a multi-objective reward function guides optimal actions for voltage regulation, harmonic suppression, unbalance mitigation, and frequency stability. Vali-dation in a MATLAB/Simulink hybrid microgrid with nonlinear loads and renewable intermittency shows that the proposed DRL controller reduces THD from 8.42% to 2.11%, VUF from 3.9% to 0.7%, and frequency deviation from 0.42 to 0.08 Hz, while improving settling time by nearly 50%. Convergence and multi-run statistical analysis further confirm the robustness, stability, and reproducibility of the trained policy, demonstrating the effectiveness of DRL as an intelligent and scalable solution for next-generation microgrid PQ control.
Microgrids play a critical role in enhancing the flexibility, reliability, and sustainability of modern power systems by integrating distributed energy resources, energy storage systems, and controllable loads. However, the inherent uncertainty of renewable generation and the stochastic nature of load demand pose significant challenges to optimal energy management. To address these issues, this paper proposes a deep reinforcement learning (DRL)-based optimal energy management framework for microgrids. The problem is formulated as a Markov decision process, where the system state captures renewable generation, load demand, and storage status, while the control actions determine power dispatch and energy storage operation. A deep reinforcement learning model is developed to learn optimal control policies through continuous interaction with the environment, enabling adaptive decision-making under dynamic and uncertain conditions. To improve learning efficiency and policy stability, state normalization and reward shaping strategies are incorporated. Furthermore, a constrained optimization mechanism is introduced to ensure operational safety and economic feasibility. Experimental results on benchmark microgrid scenarios demonstrate that the proposed method outperforms conventional rule-based strategies and model-based optimization approaches in terms of operational cost reduction, energy utilization efficiency, and robustness under uncertainty. The results indicate that the proposed DRL-based framework provides an effective and scalable solution for intelligent microgrid energy management.
Li Chen, Hong-Qiao Li, Zhen-Xing Chen et al.· European Conference on Elect...· 0 citations
The increasing integration of renewable energy sources into microgrids has intensified the need for intelligent energy management strategies capable of addressing the intermittency of solar and wind generation while ensuring reliable and cost-effective operation. Although Rule-Based Control (RBC) methods are straightforward to implement, their limited adaptability often leads to suboptimal utilisation of Hybrid Energy Storage Systems (HESS). This study develops and evaluates a Deep Reinforcement Learning (DRL)-based energy management system employing a Deep Q-Network (DQN) to coordinate battery–supercapacitor operation within a renewable microgrid. A Gymnasium-compatible simulation environment was constructed using a publicly available time-series dataset comprising renewable generation, load demand, electricity prices, battery state of charge (SoC), and supercapacitor SoC. Feature engineering, incorporating sinusoidal temporal representations and Min-Max normalisation, was applied to enhance learning stability and capture cyclical demand and generation patterns. The DQN agent was trained over 50 episodes and benchmarked against a conventional RBC strategy under identical operating conditions. Training performance demonstrated progressive policy improvement, with cumulative rewards increasing from approximately -1200 to -400, indicating enhanced decision-making capability during learning. The learned controller exhibited adaptive energy scheduling through dynamic utilisation of the supercapacitor and selective grid interaction in response to varying operating conditions, whereas the RBC followed a deterministic control strategy with limited flexibility. However, comparative evaluation revealed that the DQN did not consistently outperform the RBC in cumulative economic performance, suggesting the need for further refinement of the reward function, training process, and hyperparameter configuration. Nevertheless, the proposed framework demonstrates the feasibility of applying deep reinforcement learning to coordinated battery–supercapacitor energy management and highlights its potential to enhance operational flexibility and intelligent resource utilisation in renewable microgrids. The study contributes a dataset-driven reinforcement learning framework that provides a foundation for future research on advanced AI-based energy management systems and the integration of more sophisticated reinforcement learning algorithms for resilient and sustainable microgrid operation.
Daniel Owusu· American Journal of Neural N...· 0 citations
An intelligent Model Predictive Control framework for optimal power flow management in microgrids, with the objective of enhancing operational resilience, reducing diesel fuel consumption, and preventing blackouts through coordinated electric vehicle (EV) charging and discharging is proposed.
H. Taha, Ahmed Abdelrahman, A. Mammeri· Energy Efficiency· 0 citations
Simulation results show that the proposed Human-Interactive Lagrangian SAC (HI-LSAC) achieves significantly lower voltage-violation severity and reduced power losses compared with the other baseline methods.
An integrated deep reinforcement learning (DRL) strategy in which a single Twin Delayed Deep Deterministic Policy Gradient (TD3) agent simultaneously adjusts the proportional and integral gains of the speed controller, the torque modulation coefficient (Ks), and the regenerative braking factor (βreg) is proposed.
Saber Hadj Abdallah, Fatma Ben Salem, Jaouhar Mouine et al.· Electronics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.