Skip to content

A deep reinforcement learning framework for hybrid electric vehicle energy management: Integrating real-world data augmentation and multi-scale perception

Jul 2026 · Proceedings of the Institution of mechanical engineers. Part D, journal of automobile engineering · 0 citations · 23 references

TL;DR

A fresh enhanced training set covering all-round driving cycles based on real driving data is constructed, thus breaking through the training restrictions brought by standard driving cycles, and improves fuel economy by 5.2% on average across various driving cycles compared to adaptive ECMS (AECMS).

Abstract

To solve the generalization bottleneck and environmental perception limitation of deep reinforcement learning (DRL) in the charge-sustaining (CS) stage of hybrid electric vehicles, this paper proposes a novel adaptive hierarchical energy management strategy, which combines real-world data enhancement and multi-scale perception. Aiming at the problems of over-fitting and short-sighted decision-making commonly existing in traditional strategies, this study constructed a fresh enhanced training set covering all-round driving cycles based on real driving data, thus breaking through the training restrictions brought by standard driving cycles. In addition, the historical average speed window is introduced as the enhanced state, which enables the strategy to capture the macro traffic flow trend. In terms of control architecture, a bi-level coupled mechanism based on soft actor-critic (SAC) and equivalent consumption minimization strategy (ECMS) is designed. Experimental results indicate that across multi-type driving cycle tests, the final SOC deviation of the agent based on the augmented training set is maintained within ±2%. Under identical operating conditions, the augmented agent incorporating a 30 s average velocity observation achieves a 43.55% reduction in the standard deviation of SOC fluctuations compared to its counterpart without such observation. While ensuring battery SOC robustness, the proposed strategy improves fuel economy by 5.2% on average across various driving cycles compared to adaptive ECMS (AECMS).

View source

Similar papers

Open access Jul 2026

Integrated Deep Reinforcement Learning Framework for Adaptive PI Control and Multi-Objective Energy Management in Electric Vehicle Powertrains

An integrated deep reinforcement learning (DRL) strategy in which a single Twin Delayed Deep Deterministic Policy Gradient (TD3) agent simultaneously adjusts the proportional and integral gains of the speed controller, the torque modulation coefficient (Ks), and the regenerative braking factor (βreg) is proposed.

Saber Hadj Abdallah, Fatma Ben Salem, Jaouhar Mouine et al. · 0 citations
2026

RL-DTNet: Reinforcement Learning Driven Deep Temporal Network for Accurate State-of-Charge Estimation in Lithium-Ion Batteries

A reinforcement learning driven deep temporal network (RL-DTNet) is proposed for SoC prediction, integrating of reinforcement learning for self-correction, temporal attention to handle dynamic dependencies, and a degradation-aware model for long-term prediction accuracy.

S. M. Kanna, G. Narmadha, B. Sakthivel · 0 citations
Review Open access Jul 2026

Deep Reinforcement Learning-Based Energy and Power Management for Ships: A Perspective Review of Methods and Applications

Energy and power management systems (EMS/PMS) are essential for electric-propulsion ships, affecting propulsion performance, fuel consumption, emissions, and component lifetime. As shipboard power systems integrate heterogeneous energy resources and face nonlinearity, uncertain load demand, and multi-source interactions, deep reinforcement learning (DRL) has emerged as a promising adaptive, sequential decision-making tool in shipboard EMS/PMS. This perspective reviews DRL studies through a hierarchical decision-making framework comprising power dispatch, energy coordination, and operational strategy. Most research focuses on real-time power dispatch, while emerging research addresses energy coordination via multi-source cooperation, multi-objective operation, degradation awareness, and uncertainty handling. However, operational strategy remains underexplored, despite its role in speed control, route-aware planning, predictive operation, and voyage scheduling. This paper argues that future shipboard EMS/PMS adopt integrated hierarchical DRL frameworks across all three decision layers, leveraging DRL’s strengths in sequential policy learning in dynamic environments and supporting multi-time-scale decision-making. This paper clarifies current research trends, identifies gaps, and outlines future directions toward adaptive, reliable, and autonomous shipboard EMS/PMS in next-generation electric-propulsion ships.

Yujeong Kang, Dita Puspita, Il-Yop Chung · 0 citations
Aug 2026

Collaborative optimization of car-following control and energy management for PHEBs based on deep reinforcement learning

Simulation results demonstrate that, compared to the traditional hierarchical optimization framework, the proposed strategy achieves significant improvements in terms of mean absolute jerk, root-mean-square (RMS) value of acceleration, power demand, battery SOH degradation, battery temperature violation, and comprehensive operating cost.

Chengrui Zhang, Fei Ju, Sichen Gao et al. · 0 citations

Simplified reinforcement learning for energy management of extended-range electric vehicles based on STM32

To address the poor adaptability of traditional rule-based control, the operational instability of basic Q-learning algorithms, and the critical difficulty of deploying complex reinforcement learning models on resource-constrained on-board embedded platforms, this paper proposes a lightweight, simplified Q-learning energy management strategy for extended-range electric vehicles (REEVs), successfully implemented on an STM32 microcontroller. The algorithm achieves significant computational reduction by simplifying the traditional 5×5 state-action space into a highly condensed 2×2 grid. Furthermore, a power cooling mechanism is introduced, a multi-dimensional reward function is reconstructed to balance competing vehicle demands, and an ε-decay exploration strategy is designed. Software-in-the-loop (SIL) simulation verification demonstrates that the proposed strategy tightly controls the state-of-charge (SOC) standard deviation within 0.09. Additionally, high-frequency power fluctuations and range extender start-stop times are drastically reduced, and overall energy efficiency is improved by 50.3% compared with traditional strategies. The optimized algorithm occupies only 72.3% of RAM and 68.7% of Flash memory, fully satisfying strict on-board embedded system constraints and providing a highly feasible solution for intelligent REEV energy management.

Junyan Guo · 0 citations
Open access Aug 2026

Hierarchical Energy Management for Fuel Cell Electric Vehicles with Adaptive-Modality Deep Deterministic Policy Gradient

Fuel cell electric vehicles (FCEVs) require energy management strategies that can balance hydrogen economy, battery utilization, component protection, and real-time control under varying driving conditions. This paper proposes an Adaptive-Modality Deep Deterministic Policy Gradient and Model Predictive Control hierarchical energy management strategy (AMDDPG–MPC HEMS). In the proposed architecture, the upper-level AMDDPG controller identifies driving-condition patterns and generates adaptive weights for hydrogen consumption, battery power, and state-of-charge regulation, while the lower-level MPC controller performs constrained power allocation between the fuel cell and battery. To improve adaptability, the AMDDPG algorithm incorporates an adaptive modality perception mechanism that extracts driving-condition features and a multi-scale reward mechanism that coordinates short-term energy-saving objectives with long-term component-protection requirements. A dedicated weight-scheduling and switching mechanism is also introduced to ensure smooth transitions between operating conditions. The proposed strategy is evaluated under the World Light Vehicle Test Cycle and Urban Dynamometer Driving Schedule and compared with rule-based, equivalent consumption minimization, and fixed-weight MPC strategies. The results show that the AMDDPG–MPC HEMS achieves the lowest equivalent hydrogen consumption, with reductions of 18.853% and 11.732% relative to the rule-based strategy under the two driving cycles, respectively. It also improves fuel-cell operating efficiency and maintains feasible battery SOC regulation. These results demonstrate the effectiveness and engineering potential of the proposed hierarchical energy management strategy.

Yan-Tao Si, Zhuo Wang, Changqun Sun et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.