Skip to content

Distribution Network Optimization with Aggregation and Reinforcement Learning Under Massive Distributed Resources Integration

Aug 2026 · Energies · 0 citations · 22 references

TL;DR

A DNs cooperative optimization method based on resource aggregation and RL is proposed, which demonstrates the effectiveness and superiority of the proposed method in ensuring the secure operation of DNs.

Abstract

The integration of large-scale distributed energy resources (DERs) into distribution networks (DNs) brings challenges to the effective control of DNs. In traditional approaches, mathematical or reinforcement learning (RL)-based solution algorithms are commonly used. However, the exponential increase in the number of DERs reduces the effectiveness of these strategies. Mathematical methods struggle to cope with the dynamic uncertainty caused by the high penetration of renewable energy, while RL algorithms relying on global data training may violate multi-agent privacy protocols. This paper proposes a DNs cooperative optimization method based on resource aggregation and RL. To reduce optimization dimensionality and ensure the privacy of resource data, a dynamic aggregation strategy is employed to aggregate a large number of distributed energy resources into aggregated entities, and the adjustable active–reactive power boundaries of each aggregated entity are derived. To fully exploit the regulation capability of DNs, data centers (DCs), as novel devices, are considered as flexible loads. To improve the convergence speed of model training and decision-making accuracy, evolution strategies (ES) and prioritized experience replay (PER) are integrated into the Soft Actor-Critic (SAC) algorithm, respectively. The proposed method is validated on the IEEE 33-bus and IEEE 123-bus systems. The results demonstrate the effectiveness and superiority of the proposed method in ensuring the secure operation of DNs.

Read PDF

Similar papers

Preprint Aug 2026

Distributed coordination for transmission-distribution systems with nonlinear flexibility aggregation

High shares of distributed energy resources (DERs) transform distribution systems into active participants in integrated transmission and distribution (ITD) operations. Linear models enable scalable distribution-level flexibility aggregation but can misclassify AC feasible operating points, whereas direct nonlinear aggregation becomes costly, especially in multiperiod ITD coordination. This paper reformulates transmission-distribution coordination within a hierarchical optimization framework and introduces a non-iterative predictor-corrector aggregation method. By leveraging path-following techniques from real-time optimal control, the approach achieves tractable computation with guaranteed error bounds. Across 24 radial distribution-network cases and seven meshed variants, including the real KIT Campus North grid, the proposed method yields substantially lower sampled false- and lost-flexibility rates than linear surrogates and a convex relaxation. On two 24-period ITD testcases, the formulation reduces end-to-end wall-clock time by factors of 6 relative to the corresponding centralized formulation, primarily through dimensionality reduction.

Xinliang Dai, Yanlin Jiang, Frederik Zahn et al. · 0 citations
Open access Jul 2026

Federated Aggregation via Artificial Bee Colony and Optimal Transport for Distributed Energy Systems

A three-tier “device–edge–cloud” federated edge collaborative scheduling framework FedAOT is proposed, integrating an improved Artificial Bee Colony algorithm with Optimal Transport for adaptive aggregation optimization.

Jun Wang, Lijun Lu, Peng Li et al. · 0 citations
Open access Aug 2026

The Analysis of Multi-Scale Collaborative Optimization Scheduling for Electric Vehicle Clusters

As large-scale electric vehicle clusters become increasingly integrated into grid-wide collaborative scheduling, the cross-domain flow of massive user data introduces serious privacy risks. To address these challenges, this study proposed a distributed data privacy protection framework tailored for multiscale collaborative optimization of electric vehicle clusters. The framework mitigated single-point failures and trust issues commonly found in centralized scheduling systems. The proposed approach combined hierarchical federated learning with adaptive differential privacy to establish a three-tier collaborative architecture—vehicle, station, and cloud. At the charging station level, local models were trained with perturbed gradients, where an adaptive noise injection mechanism enforced (ε,δ)-differential privacy. At the cloud level, a multitimescale optimization model was employed: in the day-ahead stage, the globally aggregated hierarchical-federated-learning model predicted the schedulable capacity of electric vehicle clusters.

Haiyan Cao, Jianlin Qiu, Xiaodan Cai et al. · 0 citations
2026

Scalable Traffic Allocation in Dynamic Networks via End-to-End Imitation Learning

Networks with highly dynamic data transmission demands and network topologies are common in real world. A fundamental problem in such networks is achieving scalable traffic allocation to maximize long-term total throughput under link capacity constraints. However, state-of-the-art (SOTA) works lack scalability. This is primarily due to two reasons in large-scale networks: first, they require solving constrained optimization problems online, which leads to high decision latency; second, they rely on reinforcement learning algorithms for policy optimization, which are inefficient in exploration and challenging to train effectively. To address these issues, we propose the Fast Networked Control (FNC) policy framework, which firstly utilizes parallelizable neural network modules to process the state and generate raw decisions, followed by basic operations such as normalizations and comparisons, which do not require iteration or optimization, to obtain decisions that satisfy the constraints. Hence, FNC policy avoids solving constrained optimization problems and supports parallel execution, significantly reducing decision latency. Furthermore, this policy preserves gradient flow and supports backpropagation, which enable us to design an imitation learning algorithm to efficiently train the policy in an end-to-end manner. Experiments in large-scale networks show that our FNC policy achieves an average 8% improvement in demands satisfaction and 10 times reduction in decision latency versus SOTA works.

Zhaoxing Yang, Guiyun Fan, An-Jie Cao et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.