Optimal Containment of Multiagent Systems With Multistep Policy Gradient Reinforcement Learning.
This article analyzes the optimal containment control problem of discrete-time multiagent systems (MASs). Multistep temporal difference (TD) learning is integrated with policy gradient (PG) reinforcement learning (RL) to form an online off-policy multistep PG (MS-PG) algorithm. The proposed MS-PG algorithm achieves optimal control performance under completely unknown system dynamics and accommodates asynchronous policy updates among agents. The closed-loop system stability and the algorithmic convergence are rigorously established. Furthermore, an actor-critic neural network (NN) architecture is employed to approximate the control policy and the optimal Q-function, respectively, with data-driven weight update laws derived from the proposed algorithm. To improve training efficiency and sample efficiency, an experience replay (ER) mechanism is incorporated, constructing a hybrid learning framework that fully exploits both offline batch data and online operational data. Finally, simulation results verify the effectiveness of the proposed method.