Skip to content
Preprint

High-Precision Formation Control for Heterogeneous Multi-Robot Systems via Hierarchical Hybrid Physics-Informed Deep Reinforcement Learning

Jul 2026 · 0 citations · 43 references
Computer Science

TL;DR

A hierarchical hybrid physics-informed deep reinforcement learning (HHy-PIDRL) framework, aiming to realize high-precision, highly responsive formation control for heterogeneous multi-robot systems (HMRSs).

Abstract

Existing classical control methods commonly require precise models and struggle to cope with model uncertainties and external disturbances, while end-to-end reinforcement learning (RL) approaches suffer from low sample efficiency and poor convergence. To overcome these challenges, this paper proposes a hierarchical hybrid physics-informed deep reinforcement learning (HHy-PIDRL) framework, aiming to realize high-precision, highly responsive formation control for heterogeneous multi-robot systems (HMRSs). The proposed framework contains two layers. Specifically, first, the upper layer designs an autonomous navigation policy network for Ackermann-steering leader based on the Soft Actor-Critic (SAC) deep reinforcement learning (DRL) algorithm. Second, the lower module integrates a high-fidelity physical feed-forward controller, a classical proportional-derivative (PD) controller, and an adaptive DRL residual controller to propose an effective hybrid model and DRL (HM-DRL)-based formation control policy network. Third, a unique hierarchical reward function is designed for training Omnidirectional followers, which effectively guides agents toward a refined, stable control policy. Experimental results demonstrate that, the success rate of both the upper-layer autonomous navigation policy network and the HM-DRL based formation control policy networks reach 100%. Meanwhile, ablation experiments are conducted to verify the validity and credibility of the proposed method.

View source

Similar papers

Aug 2026

Deep reinforcement learning–based safe path planning for leader–follower robots

This work proposes a modified Multi-Agent Twin-Delayed Deep Deterministic Policy Gradient (M-MATD3) algorithm, specifically designed to mitigate common issues such as overestimation bias and high variance observed in standard MATD3.

Ehsan Kazemi Tameh, Mohammadreza Estarki, Saeed Khodaygan · 0 citations
Conference Jul 2026

Hybrid Computed Torque Control and Soft Actor-Critic Framework for Sample-Efficient Robotic Manipulator Control

Reinforcement Learning (RL) has emerged as a promising approach for robotic control, enabling agents to learn control policies through interaction with complex and dynamic environments. However, standalone RL methods often suffer from poor sample efficiency, limiting their practicality for real-world robotic systems. To address this limitation, recent studies have combined RL with classical controllers such as proportional–integral–derivative (PID) control, where the classical controller provides a baseline policy and RL learns residual corrective actions. Nevertheless, conventional PID controllers do not explicitly incorporate the full nonlinear manipulator dynamics.This paper proposes a physics-informed residual reinforcement learning framework that combines Computed Torque Control (CTC) with Soft Actor-Critic (SAC) for trajectory tracking of a 2-DOF robotic manipulator. The CTC component utilises analytical Lagrangian dynamics to provide a nominal control torque, while SAC learns bounded residual corrections to compensate for model uncertainties and unmodelled effects. The proposed framework is evaluated in CoppeliaSim and compared against CTC-only, RL-only, and PID+SAC baselines under identical experimental conditions.Experimental results demonstrate that the proposed CTC+SAC framework achieves the lowest mean and steady-state tracking errors among all evaluated methods, with a 4.9% reduction in mean error over RL-only and a 3.3% reduction over PID+SAC within the considered simulation setup. The results suggest that incorporating analytical robot dynamics into the residual learning framework improves tracking performance and sample efficiency compared to both pure RL and classical controller baselines.

Mona Alsbakhi, Mohammed M. Lubbad, M. Tabash et al. · 0 citations
Conference Aug 2026

Deep Reinforcement Learning-Based Intelligent Control Algorithm for Dual-Arm Robots

This paper presents a review-oriented comparative analysis of deep reinforcement learning (DRL) for intelligent control of dual-arm robots. Instead of focusing on a single control algorithm, it organizes recent studies into an algorithm-task-metric framework and extracts quantitative evidence from representative applications including cooperative grasping, assembly, transportation, obstacle-aware planning, contact-rich control, and sim-to-real transfer. PPO, MAPPO, MADDPG, and SAC are compared in terms of success rate, convergence behavior, trajectory smoothness, force regulation, safety constraints, and transferability. Key design factors and future trends, including reward design, multimodal perception, safe reinforcement learning, sample efficiency, and real-robot deployment, are summarized to provide practical guidance for dual-arm intelligent cooperative control.

Yulong Shi, Xiangxu Sun, Shengli Zhou et al. · 0 citations
Open access Jul 2026

A Belief-Driven Hybrid Reinforcement Learning Framework for Decentralized Multi-Robot Navigation Under Partial Observability

Decentralized multi-robot navigation is difficult when robots must act from local observations without centralized coordination or explicit inter-robot communication. A belief-driven hybrid reinforcement learning framework is evaluated for planar multi-robot navigation under partial observability. Each robot builds a compact local state from its position, waypoint target, sector-based proximity readings, and a decaying occupancy belief that summarizes recent obstacle evidence. A Deep Deterministic Policy Gradient (DDPG) actor produces continuous velocity proposals, and a lightweight geometric safety-blending layer combines this command with goal-seeking and reactive avoidance vectors before execution. The simulation was revised to use e-puck-compatible heading-limited forward motion rather than side-slip motion. The framework is intentionally solver-free at runtime and does not introduce online constrained optimization or new communication mechanisms. The evaluation reports a controlled five-seed study using seeds 101–105 and a 10-seed stress suite covering scalability, symmetric crossing, corridor, and dense dynamic-obstacle cases. In the controlled nominal evaluation, full three-robot completion occurred in all five runs, with 100.0% mean success, 142.2 mean steps, and no recorded collision timestep. In the hybrid stress suite, nominal, four-robot swap, five-robot crossing, symmetric-deadlock, and corridor cases achieved full success in all 10 seeds. Dense dynamic obstacles were the main failure case, with 5/10 full-success runs, 5 robot timeouts, and 10.1 mean collision events per run. These results support the feasibility of the hybrid structure in moderate tested conditions while showing that dense moving obstacles remain a practical limitation. Formal safety guarantees, matched benchmark comparisons, physical robot validation, and wider randomization remain areas requiring future work.

V. Malathi, Pramod Sreedharan, Rthuraj Puthiyaveedu Rajesh et al. · 0 citations
Open access 2026

Hierarchical Reinforcement Learning Control of a Quadruple Tank Plant Under Partial Observability

Experimental results in simulation and on real hardware demonstrate that the decentralized–supervised architecture can achieve comparable or improved aggregate tracking performance relative to a centralized policy, while preserving decentralized proposal generation and enabling execution-time supervisory coordination under partial observability.

A. Bozzi, Matteo Aicardi, E. Zero et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.