Large-scale multi-robot trajectory planning faces significant challenges in computational efficiency, scalability, trajectory quality, and safety, especially in complex, obstacle-dense environments. To address this, we propose CHORD, a hierarchical generative framework for multi-robot motion planning. At the macroscopic level, to generate efficient global guidance, we formulate the multi-robot motion as time-varying Gaussian Mixture Model (GMM) density fields. We develop a tailored Diffusion Transformer that employs structured temporal attention to capture long-range dependencies, generating coherent macroscopic distribution trajectories. To jointly optimize transport efficiency, safety, and smoothness, a cost gradient guidance mechanism integrates Wasserstein distance, Conditional Value at Risk (CVaR), and Gaussian process priors into the diffusion sampling process. At the microscopic level, a probabilistic mapping strategy is employed to generate individual reference trajectories, facilitating flexible split-and-merge behaviors. Finally, a distributed model predictive controller is developed for real-time trajectory tracking and collision avoidance. To mitigate tracking errors caused by execution uncertainties, we introduce a closed-loop replanning mechanism via generative inpainting. This allows the planner to periodically regenerate the future trajectory anchored to the current actual state of the system. Extensive simulations demonstrate CHORD’s superior scalability, computational efficiency, trajectory quality, and safety, maintaining high-quality performance with up to 500 robots while achieving a reduction of nearly three orders of magnitude in computation time compared to state-of-the-art baselines. Real-world experiments further validate its effectiveness and practical applicability. Note to Practitioners—This work addresses the computational and safety challenges of large-scale multi-robot trajectory planning. Traditional methods often fail to deliver real-time performance or high-quality trajectories when coordinating hundreds of robots in cluttered environments. CHORD introduces a hierarchical generative framework that integrates macroscopic multi-robot guidance with microscopic control. Unlike open-loop approaches, it incorporates a reactive inpainting mechanism to ensure closed-loop execution against disturbances. This ensures scalable, efficient, and safe planning, significantly reducing computation time and trajectory length. The approach is well-suited for time-critical applications like disaster response, environmental monitoring, and industrial automation.
Kang Ding, Chun-Xuan Jiao, Yun-Ze Hu et al.· IEEE Transactions on Automat...· 0 citations
Interactive navigation requires robots to actively modify cluttered environments to create traversable paths, going beyond passive obstacle avoidance. However, existing methods either depend on global maps and lack the reasoning capabilities to make interaction decisions from local observations, or are restricted to interactions with simple geometric objects, limiting their applicability in partially observable, unstructured environments. To address these challenges, we propose counterfactual interactive navigation, named CoIN, a vision-language model (VLM)-based hierarchical framework that integrates high-level interaction reasoning with low-level loco-manipulation policies for diverse objects. Specifically, we propose CoIN-VLM, a VLM that internalizes counterfactual reasoning to evaluate the effect of object removal on goal reachability, thereby deciding when interaction is necessary and which object to interact with. To further align such reasoning with the robot’s physical capabilities, we inject robot skill descriptions into the VLM context and ground them into a metric-scale environmental representation, ensuring that the generated plans remain physically feasible. To execute the generated high-level plans, we develop a comprehensive skill library through reinforcement learning, specifically introducing traversability-oriented strategies to manipulate diverse objects for path clearance. Furthermore, a systematic benchmark in Isaac Sim is proposed to evaluate both the reasoning and execution aspects of interactive navigation. Extensive simulations and real-world experiments demonstrate that CoIN significantly outperforms representative baselines, achieving a 17% higher overall success rate and over 80% improvement in complex long-horizon scenarios compared to the best-performing baseline, while exhibiting robust generalization across diverse object categories. Our project page is available at https://coins-internav.github.io/ Note to Practitioners—This work addresses the practical challenge of enabling autonomous robots to reach goals in cluttered indoor environments where the path is blocked by movable objects, without relying on a global map. The primary application is warehouse automation, facility inspection, and disaster-response robots that must decide when to interact, which object should be moved, and how to interact with diverse objects. The fine-tuned vision-language reasoning module determines the timing of interaction and selects the object whose removal is most likely to open a useful path. The learned skill library then executes efficient physical interactions, such as pushing obstacles or opening doors, to create traversable space for navigation. This improves navigation efficiency and reliability by reducing unnecessary detours and allowing the robot to complete tasks that are infeasible for passive obstacle avoidance.
Kangjie Zhou, Zhe-Jia Wen, Zhiyong Zhuo et al.· IEEE Transactions on Automat...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.