Skip to content
Open access

Integrated Deep Reinforcement Learning Framework for Adaptive PI Control and Multi-Objective Energy Management in Electric Vehicle Powertrains

Jul 2026 · Electronics · 0 citations · 33 references

TL;DR

An integrated deep reinforcement learning (DRL) strategy in which a single Twin Delayed Deep Deterministic Policy Gradient (TD3) agent simultaneously adjusts the proportional and integral gains of the speed controller, the torque modulation coefficient (Ks), and the regenerative braking factor (βreg) is proposed.

Abstract

Electric vehicle (EV) powertrains involve complex interactions between speed regulation, energy consumption, regenerative braking, and battery thermal behavior. Most existing approaches address controller tuning and energy management separately, which may limit the overall system performance. This paper proposes an integrated deep reinforcement learning (DRL) strategy in which a single Twin Delayed Deep Deterministic Policy Gradient (TD3) agent simultaneously adjusts the proportional and integral gains of the speed controller (Kpv, Kiv), the torque modulation coefficient (Ks), and the regenerative braking factor (βreg). A multi-objective reward formulation is adopted to account for speed tracking performance, energy efficiency, regenerative energy recovery, battery thermal constraints, and driving comfort. The framework is implemented through a MATLAB R2022b/Simulink–Python 3.10 co-simulation environment that enables online interaction between the EV model and the learning agent. Performance is evaluated using the Worldwide Harmonized Light Vehicle Test Procedure (WLTP). Compared with a conventional fixed-gain PI controller, the approach reduces gross energy consumption by 16.2%, decreases speed tracking error by 43.7%, increases regenerative energy recovery by 21.4%, limits battery temperature rise by 30.4%, and lowers RMS jerk by 33.7%. The results indicate that jointly optimizing control and energy management variables can improve both vehicle dynamic performance and energy utilization. The methodology offers a practical framework for the development of adaptive and intelligent control systems in future electric vehicles.

Read PDF

Similar papers

Aug 2026

Collaborative optimization of car-following control and energy management for PHEBs based on deep reinforcement learning

Simulation results demonstrate that, compared to the traditional hierarchical optimization framework, the proposed strategy achieves significant improvements in terms of mean absolute jerk, root-mean-square (RMS) value of acceleration, power demand, battery SOH degradation, battery temperature violation, and comprehensive operating cost.

Chengrui Zhang, Fei Ju, Sichen Gao et al. · 0 citations
Open access 2026

Enhanced Reinforcement Learning-Based Energy Management via Prioritized Exploration and Experience Replay for Axle-Split Hybrid Electric Vehicles

Axle-split hybrid electric vehicles (HEVs) have gained attention for their potential to improve fuel economy and enable electric all-wheel-drive operation through dual-motor (P0+P4) configurations. This configuration allows energy harvesting during deceleration events from both motors, effectively balancing braking safety with energy regeneration – a challenge in traditional P0-only systems. However, adding a P4 motor introduces additional control complexity. This study presents a reinforcement learning (RL)-based energy management strategy for a 48 V P0+P4 HEV using a Twin-Delayed Deep Deterministic Policy Gradient with Prioritized Exploration and Experience Replay (TD3-PEER) algorithm. To address the control complexity, a motor activation threshold is introduced, consolidating the on/off switching and power control of each motor into a single decision variable. Consequently, the three-power-source system requires only two control variables, one for each motor, simplifying the control task. Furthermore, this work validates the benefits of the prioritized exploration technique without relying on DP-derived expert demonstrations or supervisory control heuristics. The results reveal that the proposed TD3-PEER algorithm effectively learns a charge-sustaining strategy for a complex P0+P4 HEV system without prior knowledge of the driving cycle or DP-derived expert action guidance. In a case study, the proposed method achieves 95.1% and 94.2% of global optimality in training and validation cycles, respectively.

Yu He, K. Kwak, Kyoungseok Han et al. · 0 citations
Open access Aug 2026

Hierarchical Energy Management for Fuel Cell Electric Vehicles with Adaptive-Modality Deep Deterministic Policy Gradient

Fuel cell electric vehicles (FCEVs) require energy management strategies that can balance hydrogen economy, battery utilization, component protection, and real-time control under varying driving conditions. This paper proposes an Adaptive-Modality Deep Deterministic Policy Gradient and Model Predictive Control hierarchical energy management strategy (AMDDPG–MPC HEMS). In the proposed architecture, the upper-level AMDDPG controller identifies driving-condition patterns and generates adaptive weights for hydrogen consumption, battery power, and state-of-charge regulation, while the lower-level MPC controller performs constrained power allocation between the fuel cell and battery. To improve adaptability, the AMDDPG algorithm incorporates an adaptive modality perception mechanism that extracts driving-condition features and a multi-scale reward mechanism that coordinates short-term energy-saving objectives with long-term component-protection requirements. A dedicated weight-scheduling and switching mechanism is also introduced to ensure smooth transitions between operating conditions. The proposed strategy is evaluated under the World Light Vehicle Test Cycle and Urban Dynamometer Driving Schedule and compared with rule-based, equivalent consumption minimization, and fixed-weight MPC strategies. The results show that the AMDDPG–MPC HEMS achieves the lowest equivalent hydrogen consumption, with reductions of 18.853% and 11.732% relative to the rule-based strategy under the two driving cycles, respectively. It also improves fuel-cell operating efficiency and maintains feasible battery SOC regulation. These results demonstrate the effectiveness and engineering potential of the proposed hierarchical energy management strategy.

Yan-Tao Si, Zhuo Wang, Changqun Sun et al. · 0 citations
Open access Aug 2026

Synergistic Multi-Agent Reinforcement Learning for Energy Management in Fuel Cell Vehicles with Integrated Temperature Control

The coupled effects of power distribution and stack temperature strongly influence hydrogen economy, durability, and operating stability in proton exchange membrane fuel cell (PEMFC) vehicles. This study proposes an integrated energy–thermal management strategy based on the multi-agent deep deterministic policy gradient (MADDPG) algorithm for a PEMFC hybrid bus. The energy management agent regulates PEMFC power using vehicle demand, battery state of charge, and stack temperature, while the thermal management agent controls coolant and air mass flow rates using temperature errors and commanded PEMFC power. Under the unseen CHTC-C cycle, MADDPG reduces equivalent hydrogen consumption by 0.49% and 1.79% compared with SAC and DDPG, respectively, while remaining 2.92% above the offline dynamic programming benchmark. Under an independent real-world bus cycle, MADDPG reduces the maximum stack outlet temperature deviation by 98.56% and 98.07%relative to SAC and MPC, respectively. Additional tests under ambient temperature and aging variations show bounded thermal responses without retraining, and HIL experiments confirm real-time execution at a 1 s control period. Overall, the proposed strategy improves energy economy, thermal regulation, adaptability, and real-time applicability through coordinated power–temperature information exchange.

Pengyi Deng, Ying-Jie Ji, Jibin Yang et al. · 0 citations
Jul 2026

Deep Reinforcement Learning Framework for Adaptive Power Quality Management in Hybrid Microgrid

Convergence and multi-run statistical analysis further confirm the robustness, stability, and reproducibility of the trained policy, demonstrating the effectiveness of DRL as an intelligent and scalable solution for next-generation microgrid PQ control.

Pratibha V. Hurkadli, G. A. Kumar, T. Manjunath · 0 citations
Jul 2026

A deep reinforcement learning framework for hybrid electric vehicle energy management: Integrating real-world data augmentation and multi-scale perception

A fresh enhanced training set covering all-round driving cycles based on real driving data is constructed, thus breaking through the training restrictions brought by standard driving cycles, and improves fuel economy by 5.2% on average across various driving cycles compared to adaptive ECMS (AECMS).

Yushan Li, Lianbo Zhao, Fanyu Meng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.