Aug 2026· Intelligent Service Robotics· Vol 19· 0 citations· 41 references
TL;DR
This work proposes a modified Multi-Agent Twin-Delayed Deep Deterministic Policy Gradient (M-MATD3) algorithm, specifically designed to mitigate common issues such as overestimation bias and high variance observed in standard MATD3.
Decentralized multi-robot navigation is difficult when robots must act from local observations without centralized coordination or explicit inter-robot communication. A belief-driven hybrid reinforcement learning framework is evaluated for planar multi-robot navigation under partial observability. Each robot builds a compact local state from its position, waypoint target, sector-based proximity readings, and a decaying occupancy belief that summarizes recent obstacle evidence. A Deep Deterministic Policy Gradient (DDPG) actor produces continuous velocity proposals, and a lightweight geometric safety-blending layer combines this command with goal-seeking and reactive avoidance vectors before execution. The simulation was revised to use e-puck-compatible heading-limited forward motion rather than side-slip motion. The framework is intentionally solver-free at runtime and does not introduce online constrained optimization or new communication mechanisms. The evaluation reports a controlled five-seed study using seeds 101–105 and a 10-seed stress suite covering scalability, symmetric crossing, corridor, and dense dynamic-obstacle cases. In the controlled nominal evaluation, full three-robot completion occurred in all five runs, with 100.0% mean success, 142.2 mean steps, and no recorded collision timestep. In the hybrid stress suite, nominal, four-robot swap, five-robot crossing, symmetric-deadlock, and corridor cases achieved full success in all 10 seeds. Dense dynamic obstacles were the main failure case, with 5/10 full-success runs, 5 robot timeouts, and 10.1 mean collision events per run. These results support the feasibility of the hybrid structure in moderate tested conditions while showing that dense moving obstacles remain a practical limitation. Formal safety guarantees, matched benchmark comparisons, physical robot validation, and wider randomization remain areas requiring future work.
V. Malathi, Pramod Sreedharan, Rthuraj Puthiyaveedu Rajesh et al.· Robotics· 0 citations
Reinforcement learning (RL) has shown considerable promise for robotic decision-making, yet deploying multi-agent RL (MARL) on physical multi-robot systems in industrial environments remains challenging. This paper investigates the real-world applicability of decentralized MARL for multi-robot multi-machine tending. We propose Feature-fusion Multi-Agent Proximal Policy Optimization (FMAPPO), which fuses 2D LiDAR measurements with task-specific state information to enable safe decentralized multi-robot task assignment and navigation. A complete simulation-to-reality pipeline was developed using high-fidelity robotic simulation and ROS2 and deployed on physical mobile-manipulator platforms operating under realistic real-world conditions, with the robotic arms disabled during the experiments. We further investigate the sensitivity of the learned policy to command update frequency, an important consideration for real-world deployment. Comparative evaluation in simulation demonstrated that FMAPPO significantly outperformed state-of-the-art baselines with a large effect size, achieving improvements of 106\% and 21\% in parts delivery and 48\% and 11\% in parts collection over MAPPO and SMAPPO, respectively. FMAPPO also increased machine utilization by 31 and 10 percentage points, respectively, while reducing collisions by 18\% and 15\% and increasing the safety score by 14 and 6 percentage points compared with MAPPO and SMAPPO, respectively. Furthermore, real-world experiments demonstrated that the learned decentralized policies can coordinate multiple robots to service multiple machines while maintaining safe operation under real-world sensing and control constraints. Videos of the real-world experiment are available online https://anonymouspapers123.github.io/FMAPPO/.
A. Abdalwhab, Giovanni Beltrame, David St-Onge· 0 citations
Real-time obstacle avoidance is a challenge in mobile robotics, as it is an ongoing process and remains difficult to achieve in crowded, dynamic environments, where conventional planning algorithms, such as local planners, often offer limited adaptability. This paper presents a Proximal Policy Optimization-based deep reinforcement learning approach for real-time obstacle avoidance for mobile robots. The proposed system is end-to-end policy learning based on inputs from LiDAR and other auxiliary sensors, and is trained in a Gazebo-ROS environment using domain randomization to enhance robustness to sim-to-real transfer. The framework was implemented on a TurtleBot3 Burger platform and tested both in simulation and in an indoor physical environment with varying numbers of obstacles. In simulation, the proposed policy achieved a success rate of 94.2%, a 68.6% reduction in collision rate compared to the Dynamic Window Approach baseline policy, a path efficiency of 16.5%, and a 14.6% reduction in average time to goal. In real experiments, the policy has maintained success rates above 88, even under high-density conditions. The optimized onboard inference pipeline achieved less than 20 ms latency and over 50 Hz throughput on embedded hardware. These results indicate that the proposed framework is a successful and computationally feasible solution to real-time robotic navigation in dynamic environments.
Roja Ba, Priyanka Mishra, M. Kalaimani et al.· Future Technology· 0 citations
This work introduces a unified RL formulation that jointly optimizes agent and environment policies, where the environment policy learns graph edge costs to provide global movement guidance via backward Dijkstra search and achieves significant improvements over the strong search-based planner, Causal-PIBT, across multiple high-density maps.
He Jiang, Jingtian Yan, Yulun Zhang et al.· 0 citations
Extensive simulations and real-world experiments demonstrate that the proposed framework can efficiently generate and iteratively improve motion plans for different planning objectives, robotic platforms, and swarm configurations, highlighting its effectiveness, computational efficiency, and scalability as a general planning methodology.
Shuli Lv, Pengda Mao, Chen Min et al.· 0 citations
This follow-up work tests the feasibility of the neuro-inspired self-supervised learning framework for trajectory planning that leverages forward and inverse models as the internal supervisory mechanism in an environment that contains an obstacle, and demonstrates the tendency of the planner to exploit the learning signal provided by the forward and inverse models.
M. Krupa, Miroslav Cibula, Kristína Malinovská· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.