Aug 2026· Nonlinear dynamics· Vol 114· 0 citations· 42 references
TL;DR
A dynamic event-triggered mechanism (DETM) is constructed to reduce redundant controller-to-actuator signal transmissions and weight update laws are derived from the negative gradients of positive definite functions associated with the Hamilton-Jacobi-Bellman (HJB) equation.
This work proposes a reinforcement learning (RL) based optimal distributed control algorithm for the multi-agent systems (MASs) with stochastic uncertainties that uses the actor-critic-identifier structure and provides a Lyapunov-based stability proof that guarantees all errors are bounded, ensuring precise tracking between the leader and followers.
Ziming Wang, Bingbing Li, Karl H. Johansson et al.· arXiv.org· 0 citations
This study discusses the optimal consensus issue for heterogeneous multi-agent systems (MASs) defined by partially unknown dynamics within a graphical game framework. Data-driven reinforcement learning has shown efficacy in such systems; however, conventional implementations frequently depend on continuous time data transmission, which puts too much strain on computers and communication systems. This paper proposes a new event-triggered, data-based reinforcement learning control scheme to fix these problems. By adding an event-triggered mechanism (ETM) to the heterogeneous graphical game formulation, the control protocol is only updated when certain error thresholds are crossed. This uses much less resources than time-triggered methods. An off-policy integral reinforcement learning (IRL) algorithm is devised to ascertain the Nash equilibrium solution utilizing quantifiable system data, thereby obviating the necessity for precise knowledge of the internal system matrices. A theoretical analysis using Lyapunov stability theory shows that the suggested event-triggered strategy ensures asymptotic consensus and convergence to the best control weights. Finally, numerical simulation examples show that the proposed approach works well and is better than other methods. These examples show that the controller update frequency and communication load can be greatly reduced while keeping the system stable and optimal.
This article analyzes the optimal containment control problem of discrete-time multiagent systems (MASs). Multistep temporal difference (TD) learning is integrated with policy gradient (PG) reinforcement learning (RL) to form an online off-policy multistep PG (MS-PG) algorithm. The proposed MS-PG algorithm achieves optimal control performance under completely unknown system dynamics and accommodates asynchronous policy updates among agents. The closed-loop system stability and the algorithmic convergence are rigorously established. Furthermore, an actor-critic neural network (NN) architecture is employed to approximate the control policy and the optimal Q-function, respectively, with data-driven weight update laws derived from the proposed algorithm. To improve training efficiency and sample efficiency, an experience replay (ER) mechanism is incorporated, constructing a hybrid learning framework that fully exploits both offline batch data and online operational data. Finally, simulation results verify the effectiveness of the proposed method.
Kaitian Chen, Huaicheng Yan, Qiwei Liu et al.· IEEE Transactions on Cyberne...· 0 citations
This article addresses the optimal stabilization of unknown nonlinear networked industrial systems (NISs) with limited communication resources, proposing an innovative self-triggered approximate optimal control framework fused with generalized fuzzy hyperbolic model (GFHM) and reinforcement learning (RL) gradient descent. To handle unknown dynamics without prior model information, a GFHM-based state identifier is constructed, leveraging its universal approximation, flexible structure, and fewer parameters to adaptively reconstruct nonlinear dynamics. A self-triggered mechanism predicts the next control update instant via software, eliminating continuous hardware monitoring and reducing online computing cost, network data transmission overhead, and actuator energy usage. Within the RL framework, a critic neural network (NN) is established, with weights tuned via gradient descent to approximate the Hamilton-Jacobi-Bellman (HJB) equation solution and derive the approximate optimal control policy. Theoretical analysis proves the closed-loop signals are uniformly ultimately bounded (UUB). Numerical simulations validate the scheme's effectiveness in ensuring optimal control performance while saving resources.
Jian Liu, Jingjing Xia, Lijuan Zha et al.· IEEE Transactions on Industr...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.