Results validate the effectiveness of bilevel RL for complex energy optimization problems, highlighting its potential as a scalable control paradigm for smart management systems.
Abstract
This paper presents a bilevel Reinforcement Learning (RL) framework for optimizing Electric Vehicle (EV) charging through price-mediated coordination between grid operators and charging stations. Unlike prior work relying on direct control or manual subgoal engineering, the proposed approach uses dynamic pricing as an implicit coordination signal to address a complex multi-objective optimization problem involving grid stability, user satisfaction, and economic efficiency. To manage this complexity, the problem is decomposed into two levels comprised of an upper-level Distribution System Operator (DSO) that determines dynamic pricing strategies, and multiple lower-level Load Aggregators (LAs) responsible for EV charging decisions at individual stations in response to these prices. This bilevel structure captures the leader–follower interaction between DSOs and LAs, with each level operating at different temporal scales. Deep Deterministic Policy Gradient (DDPG) agents are deployed at both levels, enabling adaptive decision-making under operational constraints. Extensive simulations compare the framework against multiple Rule-Based Control (RBC) baselines. Results demonstrate that the DDPG-based DSO achieves a 42.4% higher mean reward and 19.1% higher profit compared to the best-performing RBC baseline, while preserving grid stability and user satisfaction. These results validate the effectiveness of bilevel RL for complex energy optimization problems, highlighting its potential as a scalable control paradigm for smart management systems.
A collaborative optimization framework based on multi-agent reinforcement learning is proposed for orderly charging at electric vehicle charging stations and coordinated interaction with the power grid, providing a technical reference for intelligent charging coordination under grid interaction and electromagnetic compatibility constraints.
Evaluation of this approach using a state-of-the-art commercial solver with stochastic EV rental requests under different confidence levels and time-varying electricity prices demonstrates significant benefits of the integrated mechanism design, including a reduction in charging costs and battery capacity degradation compared to the prevailing business-as-usual (BAU) approach and state of the art Laxity-based charging (LC) approach.
A hybrid optimization framework that combines greedy initialization with reinforcement learning to efficiently explore the charging station deployment problem is proposed and demonstrates stable performance across three evaluated deployment scenarios, indicating its potential applicability to increasingly complex charging infrastructure planning problems.
The increasing penetration of distributed renewable energy sources has intensified the need for intelligent bidding strategies in virtual power plants (VPPs), where reliable communication and real-time information exchange are essential for coordinated energy management. This study proposes an optimal bidding path construction framework based on the Deep Deterministic Policy Gradient (DDPG) reinforcement learning algorithm for VPP participation in electricity spot markets. A Markov decision process is established to characterize dynamic market interactions, and customized state-space optimization, constrained action-space design, and a multi-objective reward function are integrated into the Actor–Critic architecture to jointly maximize economic returns while satisfying operational constraints. The framework further incorporates communication-aware resource coordination mechanisms that leverage edge computing and low-latency information exchange to enhance decision consistency under uncertain renewable generation and volatile market conditions. Experimental evaluation demonstrates that the improved DDPG algorithm increases average daily revenue by 39.1% compared with conventional DDPG, accelerates convergence by approximately 15%, reduces revenue volatility by 12%, and maintains the constraint violation rate at 1.2%. In addition to intelligent energy scheduling, the proposed methodology provides valuable insights into communicationenabled power systems, distributed electromagnetic information networks, and wireless coordination infrastructures requiring adaptive decision-making and reliable multi-node information
P. Hu, N. Shen, H. Guo et al.· Advanced Electromagnetics· 0 citations
Electric Vehicles (EVs) are emerging as sustainable alternatives to internal combustion engine vehicles; however, efficient route planning remains a major challenge due to limited driving range, sparse charging infrastructure, and variable energy consumption patterns. Traditional shortest-path algorithms, such as Dijkstra’s and A*, often fail to account for EV-specific factors, including charging station availability, connector compatibility, and energy constraints. This study presents a comprehensive EV route optimization framework that integrates reinforcement learning (RL) with graph-based methods. A novel Dual Q–Adaptive Weighting model that balances reward and cost through a primal–dual learning mechanism is proposed. The framework learns energy-aware routing strategies from historical navigation experience. The model is compared against standard RL approaches—Q-Learning and Double Q-Learning—as well as enhanced variants of A* and Dijkstra’s algorithms that incorporate charging density and time-penalty considerations. Real-world EV charging infrastructure data from the Alternative Fuels Data Center (AFDC) and Placekey datasets are used to construct a clustered navigation graph via DBSCAN. Experimental results across multiple intercity routes show that the proposed Dual Q–Adaptive model achieves the highest route accuracy of 78.66%, outperforming Double Q-Learning (76.27%), Q-Learning (77.52%), and traditional A* (74.26%) and Dijkstra (60.92%) algorithms. A* and Dijkstra with modifications, use fewer charging stops than traditional algorithms. The Improvised algorithms provide substantial improvements over their baseline counterparts. The results demonstrate that reinforcement learning integrated with graph-theoretic optimization can enable scalable, infrastructure-aware, and efficient EV route planning.
Sarvesh Kumar, Rayappa David Amar Raj, Archana Pallakonda et al.· Scientific Reports· 0 citations
With the large-scale integration of Electric Vehicles (EVs) into distribution systems, the spatiotemporal uncertainty of charging loads and the interplay between user charging behavior and network operational constraints present new challenges to the safe and economical operation of the power system. To address the insufficient coordination among flexibility characterization, distributed optimization, and user-side responses, this paper proposes a closed-loop collaborative dispatch strategy. The strategy integrates flexibility aggregation, endogenous dynamic pricing, and user charging-station selection behavior. Firstly, a three-tier collaborative architecture comprising the Distribution System Operator (DSO), Electric Vehicle Aggregators (EVAs), and EV users is established, with rolling updates implemented using Model Predictive Control (MPC). A flexible aggregation model is developed based on set operations of vehicle-level constraints, dynamically calculating power boundaries and energy feasibility domains. Furthermore, a distributed coordinated optimization model between the DSO and multiple EVAs is established and solved via the Alternating Direction Method of Multipliers (ADMM) under privacy-preserving conditions. By analyzing the correlation between ADMM dual variables and the marginal value of network constraints, a Distribution Locational Marginal Pricing (DLMP) -inspired dynamic price signal—endogenous to the optimization—is constructed to guide spatial reallocation of charging loads. Joint simulations based on an IEEE 33-node distribution network and the Sioux Falls transport network demonstrate that the proposed strategy reduces 24 h network losses from 9.05 MWh (uncoordinated) to 8.41 MWh, lowers user total costs from 22,200 yuan to 7100 yuan, and eliminates voltage limit violations (duration reduced from 1.50 h to 0), while exhibiting good distributed solution performance and closed-loop control capability.
Si-Zu Hou, Yao Sang, Xuan Zhao et al.· Energies· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.