To address the challenge that single algorithms struggle to balance global exploration and local obstacle avoidance, and are prone to falling into local optima in complex environments, this paper proposes a Strategic Hierarchical Path Planning (SHPP) framework. This framework decouples the 3D navigation task into three synergistic layers. The top layer employs a reinforcement learning network equipped with a Credit Alignment Mechanism (CAM) to provide macroscopic guidance, eliminating credit assignment pollution and escaping local minima. The middle layer introduces a Dual-Guided Particle Swarm Optimization (DG-PSO) algorithm to map discrete commands into continuous smooth trajectories. The bottom layer executes physical collision avoidance and tracking based on the Artificial Potential Field (APF) method. Simulations indicate that the system can establish a stable policy in approximately 625 episodes, achieving an average reward of 95.6. Furthermore, benefiting from the hierarchical architecture's smooth optimization in continuous space, the average flight path length is 138.3 meters, a reduction of approximately 19% compared to traditional discrete decision-making models. These quantitative results fully validate the superior performance of the proposed architecture in complex 3D environments.
Fei Wang, Jun-Yong Shi, Zhao-Kun Chen et al.· International Conference on...· 0 citations
To address the issues of extreme time-consuming bottlenecks such as special heat treatments in the manufacturing of aerospace precision forgings and scheduling failures caused by sudden rush orders, an intelligent scheduling model based on Graph Neural Networks (GNN) and Deep Reinforcement Learning (DRL) is proposed. The model uses disjunctive graphs to rigorously represent the topology of forgings processes and equipment status, achieving Markov Decision Process modeling for dynamic production line scheduling; a scale-independent policy network is designed to extract highdimensional features, and the Proximal Policy Optimization (PPO) algorithm is employed for autonomous training. Experiments show that under harsh conditions such as dynamic rush orders and extreme bottleneck surges, the model's solution quality and robustness are significantly superior to heuristic rules like SPT and MWKR.Under the zero - shot condition, the small - scale (15×15) trained policy can be directly generalized to ultra - large - scale (30×20) scenarios. The single - instance inference only takes 2.7 seconds, providing a new paradigm for the hard real - time scheduling of aerospace precision forgings.
Kai-Kai Liu, Zheng-Xiang Ma, Wei-Chao Yu et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.