Action-Space-Oriented Reinforcement Learning Compensation for PI-TPS-Controlled Low-Voltage DAB Converters Under Input-Voltage and Load Variations
This study examines how the location of reinforcement-learning (RL) compensation affects the control of a low-voltage dual-active-bridge converter using triple-phase-shift modulation. A proportional–integral triple-phase-shift controller is used as the baseline, and two deep deterministic policy gradient schemes are compared. The first directly corrects the three modulation variables, whereas the second adjusts the phase-shift command before it is mapped to those variables. The three control structures are tested at the rated operating point, during fixed and stepwise input-voltage changes from 80 to 120 V, and under fixed and stepwise load changes at an input voltage of 100 V. Direct correction of the modulation variables gives no consistent improvement, mainly because the variables are strongly coupled. The phase-shift-level scheme avoids this difficulty by leaving the modulation mapping to the existing controller. At the rated point, it reduces the steady-state voltage error by 29.34% and the full-transient peak inductor current from 69.82 to 36.84 A. During the load tests, it also gives lower voltage ripple and peak inductor current and performs better during load transitions. For this converter, placing the RL action at the higher phase-shift level is more effective than directly modifying the coupled modulation variables.