Skip to content

Hierarchical Proximal Policy Optimization for Wide-Area Damping Control of Inter-Area Oscillations in Power Systems with DFIG Wind Integration

Jul 2026 · 2026 IEEE Canadian Atlantic Ocean Symposium (CAOS) · pp. 123-128 · 0 citations · 19 references

Abstract

This paper presents a hierarchical reinforcement learning framework for wide-area damping control (WADC) of inter-area oscillations in the IEEE 39-bus power system with DFIG wind farm integration driven by real variable wind speed data recorded in the Ottawa region (June 2025). A two-level architecture is developed: a base-level Proximal Policy Optimization (PPO) agent learns damping control policies from a high-fidelity Simulink model, while a meta-level controller adaptively tunes PPO hyperparameters between training batches using three complementary feedback signals (immediate, moving-window, and cumulative-historical deltas). The agent observes six PMU signals and outputs a single continuous DFIG reactive power modulation command. Training spans 35 batches of 20 episodes (700 total, ≈121 million timesteps). Evaluation against a conventional PSS baseline demonstrates 62–88% MSE reduction in generator active power oscillations across all three areas, 49% peak deviation reduction at the DFIG bus, 58.5% voltage sag reduction at the PCC, 46.5% frequency nadir improvement, and 60.8% peak angular separation reduction between areas, confirming that a DFIG wind farm replacing conventional synchronous generation can simultaneously serve as an effective wide-area damping controller.

View source

Similar papers

Open access 2026

Bi-Timescale Hierarchical Safe Reinforcement Learning for Coordinated Virtual Inertia and Damping Control in Grid-Forming Converters

The increasing penetration of inverter-based renewable generation has reduced system inertia and posed new challenges to frequency stability in modern power systems. Virtual synchronous generator (VSG) control can provide virtual inertia and damping support, but its performance strongly depends on the proper coordination of these parameters under varying operating conditions. Existing adaptive and reinforcement-learning-based methods usually regulate virtual inertia and damping at the same timescale, which may ignore their distinct physical roles and lead to coupled parameter variations. To address this issue, this paper proposes a bi-timescale hierarchical safe reinforcement learning framework, termed BiTS-HSRL-JD, for coordinated virtual inertia and damping control in grid-forming converters. In the proposed framework, a slow-timescale policy schedules virtual inertia according to operating conditions, while a fast-timescale policy adjusts the damping coefficient to suppress transient oscillations. A stage-aware state representation and a safety projection layer are further introduced to improve transient adaptability and enforce practical constraints on parameter bounds and variation rates. The proposed method is validated using a MATLAB/Simulink-based VSG system under strong-grid and weak-grid conditions, power-step disturbances, and load-switching events. Comparative results show that BiTS-HSRL-JD reduces RoCoF, improves frequency recovery, and suppresses oscillations more effectively than fixed-parameter, rule-based adaptive, and single-policy reinforcement learning methods.

Zhilin Dong, Haoqing Xiong, Rongqian Su et al. · 0 citations
Open access Aug 2026

GA–SQP Hybrid Optimization Control Strategy for Hydropower Units Oriented to Multiple Operating Conditions Under Isolated Grid Mode

Hydropower units operating in isolated grids are characterized by low rotational inertia and weak damping, making it difficult to balance rapid frequency regulation and overshoot suppression. To address this issue, this paper proposes a GA–SQP hybrid optimization control strategy for multiple operating conditions based on a high-fidelity nonlinear dynamic model. Deep feedforward neural networks are first employed to reconstruct the nonlinear torque and discharge characteristics of the hydro-turbine, providing smooth and continuously differentiable mappings for subsequent gradient-based optimization. An improved performance index combining the Integral of Time-Cubed Absolute Error (ITCAE) with a transient overshoot penalty is then formulated to suppress long-tail errors and prioritize smooth responses with reduced transient overshoot. A two-stage optimization framework is further developed, in which the Genetic Algorithm (GA) performs global exploration to identify a promising parameter region, followed by Sequential Quadratic Programming (SQP) for high-precision local refinement. Comparative simulations under low-, rated-, and high-head high-load conditions show that the proposed strategy achieves higher optimization accuracy with fewer iterative resources. Within the investigated operating range, the optimized controller maintains a very low overshoot level while preserving satisfactory response speed, effectively improving the balance between rapidity and stability in isolated-grid frequency regulation.

Fanglin Wang, Feng Gu, Ke Kang et al. · 0 citations
Open access Jul 2026

Intelligent power flow control of AC/DC hybrid transmission corridors using safe reinforcement learning agents.

As renewable generation progressively displaces conventional generators, power flow through geographically constrained transmission corridors increasingly approaches or violates thermal and stability limits, exposing the grid to congestion-induced renewable curtailment and cascading-failure risks. Traditional real-time dispatch practices, which rely on precomputed look-up tables and operator heuristics, prove inadequate when faced with rapidly growing uncertainties arising from high penetrations of wind and photovoltaic generation. This paper presents a safe reinforcement learning (SRL)-driven coordinated control framework that simultaneously regulates embedded HVDC links and dispatchable generators to enhance the transfer capability of AC/DC hybrid transmission corridors. A perturbation-based sensitivity approach distills the generator fleet into a compact subset whose output variations most strongly affect the transmission corridor power flow, effectively compressing the decision dimensionality. The sequential decision task is formulated as a Markov Decision Process model, where SRL agents are trained to govern HVDC flow and generator redispatch, under a maximum-entropy actor-critic framework, yielding policies that are simultaneously exploratory, reward-seeking, and constraint-respecting. Extensive simulation experiments and commissioning on the Yangtze River-crossing transmission corridor confirm that the SRL agent's policies elevate the mean aggregate transfer by 629 MW and raise the delivery ceiling by 807 MW, peaking at 2761 MW in heavily stressed scenarios.

Haifeng Li, Zhiwei Wang, Tao Jin et al. · 0 citations
Open access Jul 2026

Reinforcement Learning-Based Adaptive Control for a Permanent Magnet Synchronous Generator Connected to a Hybrid AC/DC Grid with Virtual Inertia Support

The increasing penetration of renewable energy sources has increased the need for advanced control strategies capable of maintaining stability under low-inertia, converter-dominated operating conditions. In grid-connected wind energy conversion systems (WECSs), constant power loads (CPLs) exhibit negative incremental impedance characteristics that can amplify DC-link oscillations and complicate the coordination between the electrical and mechanical subsystems. The main contribution of this work is a Soft Actor–Critic (SAC) reinforcement learning algorithm that tunes the outer proportional-integral gains of the machine-side DC-voltage-squared control loop together with the active damping gain, allowing online adaptation of the controller according to the operating condition and disturbance level, thereby improving energy system sustainability. The proposed control framework includes a two-mass shaft model, virtual inertia control, and DC-link load uncertainty in the form of both resistive loads and CPLs. The system is modeled and evaluated using MATLAB/Simulink, and its performance is compared with that of a conventional fixed-gain controller under AC load disturbances and wind speed variations. It has been found that for a 25% load disturbance, the maximum DC-link voltage deviation is reduced by 1.2% under resistive loading and 6.5% under CPL operation. For a 1 m/s reduction in wind speed, the corresponding reductions are 0.8% and 0.9%, respectively. The proposed controller also provides smoother output power and improved damping of the rotor speed and system frequency responses.

Islam A. Zenhom, M. Marei, Ahmed M. I. Mohamad · 0 citations
Open access Aug 2026

Comparative Performance Evaluation of Feedback Error Learning and Adaptive Neuron-PI Controllers for Small-Signal Stability Enhancement in a Single-Machine Infinite-Bus Power System

Small-signal stability remains a fundamental challenge in the operation of modern power systems, particularly as grid complexity increases with the integration of highly dynamic power sources. Conventional Power System Stabilizers (CPSSs) have been widely adopted to improve damping performance; however, their effectiveness deteriorates when the operating conditions deviate from the design point. Consequently, intelligent adaptive control techniques have attracted significant attention owing to their ability to adjust controller parameters online in response to system dynamics. This paper presents a comprehensive comparative investigation of two intelligent adaptive control strategies, namely the Feedback Error Learning (FEL) controller and the Adaptive Neuron-Proportional-Integral (Neuron-PI) controller, for enhancing the small-signal stability of a Single-Machine Infinite-Bus (SMIB) power system. Unlike previous studies that investigated each controller independently, the proposed work evaluates both controllers under identical operating conditions, using the same SMIB model, disturbances, and simulation parameters to ensure a fair and objective comparison. The performance of both controllers is evaluated using MATLAB-SIMULINK 2025 software for multiple operating conditions. The simulation results demonstrate that the FEL controller exhibits superior adaptive learning characteristics and excellent robustness under varying operating conditions.

Alyaseh Askir · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.