A role-of-learning taxonomy is proposed that categorizes existing methods according to how learning participates in the planning pipeline, including direct policy learning, learning-augmented classical planning, hybrid planning, and training enhancement methods.
Abstract
Motion planning in dynamic environments is a fundamental problem in robotics, aiming to generate safe and efficient paths, trajectories, or control actions in the presence of moving obstacles, uncertain predictions, and multi-agent interactions. It has broad applications in autonomous driving, service robotics, warehouse logistics, human-robot collaboration, crowd navigation, and multi-robot systems. This survey reviews representative works published primarily between 2015 and 2025, with a particular focus on how recent learning-based advances extend, complement, or interact with classical planning foundations. We first revisit classical planning methods as algorithmic foundations and reference frameworks for learning-based extensions. We then propose a role-of-learning taxonomy that categorizes existing methods according to how learning participates in the planning pipeline, including direct policy learning, learning-augmented classical planning, hybrid planning, and training enhancement methods. For each category, we summarize the main problem settings, representative algorithms, key ideas, integration mechanisms, strengths, and limitations. We further analyze how observation representations, prediction uncertainty, interaction modeling, planner integration, safety constraints, and training strategies shape learning-based motion planning in dynamic environments. Finally, we discuss open challenges and future directions, including sim-to-real gap, safe and certifiable planning, dense crowd navigation, perception-planning coupling, and embodied AI.
Extensive simulations and real-world experiments demonstrate that the proposed framework can efficiently generate and iteratively improve motion plans for different planning objectives, robotic platforms, and swarm configurations, highlighting its effectiveness, computational efficiency, and scalability as a general planning methodology.
Shuli Lv, Pengda Mao, Chen Min et al.· 0 citations
Intelligent robot navigation in dynamic environments remains one of the most challenging problems in autonomous robotics because navigation systems must continuously perceive environmental changes, predict moving obstacles, and generate safe trajectories while maintaining operational efficiency. Traditional navigation approaches, including graph-based path planning, rule-based obstacle avoidance, and probabilistic localization, often exhibit limited adaptability when environmental conditions change rapidly. Recent advances in artificial intelligence, particularly Deep Reinforcement Learning (DRL), have enabled autonomous robots to learn navigation policies directly from environmental interactions without relying exclusively on handcrafted rules. DRL integrates perception, decision-making, and continuous learning into a unified framework, making it particularly suitable for complex and uncertain environments such as warehouses, hospitals, manufacturing plants, urban streets, and disaster-response scenarios.
This research-review paper presents a comprehensive analysis of Deep Reinforcement Learning-based intelligent robot navigation with emphasis on dynamic obstacle avoidance, adaptive path planning, perception integration, reward optimization, and policy learning. The paper synthesizes contemporary studies related to artificial intelligence, semantic decision intelligence, cyber-physical systems, cloud intelligence, autonomous optimization, and intelligent infrastructure to establish a multidisciplinary understanding of modern robotic navigation. Particular attention is devoted to semantic AI-enabled decision intelligence, which enhances contextual understanding during navigation and improves policy robustness in continuously evolving environments (Goyal, 2025).
A conceptual DRL navigation framework is proposed comprising environmental perception, state representation, policy optimization, experience replay, reward engineering, semantic reasoning, and adaptive trajectory generation. The framework demonstrates how semantic knowledge, sensor fusion, and reinforcement learning cooperate to produce robust navigation strategies under uncertainty. Furthermore, the paper evaluates challenges involving sparse rewards, safety constraints, computational complexity, sim-to-real transfer, multi-agent coordination, and real-time deployment.
Hiroshi Tanaka· European International Journ...· 0 citations
This work proposes a modified Multi-Agent Twin-Delayed Deep Deterministic Policy Gradient (M-MATD3) algorithm, specifically designed to mitigate common issues such as overestimation bias and high variance observed in standard MATD3.
It is argued that the future of robot path planning will be dominated by hybrid systems that combine global planning, local replanning, optimization, and learning-based prediction, enabling robots to operate more safely, intelligently, and adaptively in complex real-world environments.
Chanyu Wang· Theoretical and Natural Scie...· 0 citations
This follow-up work tests the feasibility of the neuro-inspired self-supervised learning framework for trajectory planning that leverages forward and inverse models as the internal supervisory mechanism in an environment that contains an obstacle, and demonstrates the tendency of the planner to exploit the learning signal provided by the forward and inverse models.
M. Krupa, Miroslav Cibula, Kristína Malinovská· arXiv.org· 0 citations
: This paper provides a thorough survey and integrative presentation of cooperative path planning for multi-robot systems operating in dynamic, cluttered, and partially observable environments. People synthesise algorithmic foundations ranging from heuristic graph search to sampling-based motion planners, including A*, D* Lite, and Safe Interval Path Planning for discrete/time-augmented spaces, as well as RRT, RRT*, and Informed RRT* for continuous configuration spaces. Multi-agent coordination techniques are reviewed, covering reciprocal collision avoidance (ORCA) and centralised Multi-Agent Path Finding (MAPF) solvers such as Conflict-Based Search (CBS) and bounded-suboptimal variants (ECBS). The paper also examine control and safety layers like Model Predictive Control and Control Barrier Functions that translate plans into dynamically feasible commands with safety guarantees. Recent progress in cooperative multi-agent reinforcement learning (MAPPO, QMIX, MADDPG) is evaluated for adaptability under partial observability and nonstationary environments. Applications in warehousing, intelligent transportation, and disaster response are used to illustrate practical trade-offs and integration patterns, referencing real-world systems such as Kiva-style warehouse fleets and autonomous driving pipelines. The paper concludes with a focused discussion on open challenges — scalability with guarantees, safety under uncertainty, sim-to-real transfer, and planning – control interface fragility — and proposes research directions including learning-augmented heuristics, unified safety-aware planning, adaptive MPC – CBF filters, and more informative benchmarks to drive reproducible progress.
Yun Pan· Proceedings of the 3rd Inter...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.