Deep Reinforcement Learning with State Uncertainty Estimation for Continuous Control of Dynamic Complex Systems
Deep reinforcement learning has achieved remarkable success in continuous control tasks. Modeling and decision-making for dynamic complex systems remain particularly challenging due to nonlinear dynamics, stochastic disturbances, and uncertain observations, making uncertainty-aware learning an increasingly important research direction. However, policy learning often becomes unstable when environmental states are partially observable or affected by uncertainty, leading to inaccurate decision-making and degraded control performance. To address this challenge, this paper proposes a Deep Reinforcement Learning framework for Continuous Control with State Uncertainty Estimation based on the Multi-Agent Policy Collaborative Optimization (MAPCO) mechanism. The proposed framework incorporates state uncertainty estimation into policy learning while integrating parameter consensus, representation consensus, and action consensus within a unified Lagrangian optimization framework, enabling robust state representation and stable policy optimization through distributed information exchange. Extensive experiments on Cooperative Navigation and Distributed Energy Management tasks demonstrate that the proposed method consistently outperforms representative baseline approaches in cumulative reward, convergence efficiency, robustness, and control stability under uncertain environments. The results verify that the proposed framework effectively enhances continuous control performance in the presence of state uncertainty while providing an efficient and reliable solution for intelligent automated control systems.