Learning from Hindsight is presented, which brings hindsight relabeling to RL post-training of VLAs by scoring failed rollouts against the tasks they actually achieved, and achieves 5 times improvement in sample efficiency, and outperforms a dense progress-reward baseline on out-of-distribution LIBERO-PRO tasks.
Iris Xu, Sunshine Jiang, John Marangola et al.· arXiv.org· 0 citations
Mobile manipulators on construction sites offer considerable potential for increasing productivity, as the transport of materials and the execution of precise assembly work can be increasingly automated. However, the coordination of several such robots is a complex planning task, as task assignment, navigation, and reachability planning must be solved simultaneously and under dynamic environmental conditions. This work presents a multi-agent reinforcement learning (RL) approach that enables multiple mobile manipulators to complete a set of tasks in a structured environment. Each agent makes decentralized decisions about task selection and navigation, with kinematic reachability ensured by an integrated inverse kinematic solver. The policy is trained using proximal policy optimization (PPO), supported by a reward function that encourages both navigation progress and efficient task distribution. Simulation results show that the trained model is able to efficiently distribute tasks among multiple robots while taking kinematic constraints into account. The proposed method is superior to a greedy baseline that selects the nearest available task. With four robots and 35 tasks, the multi-agent RL approach achieves a success rate of 100%, while the baseline reaches only 55%.
Charlotte Stein, Yuheng Zhi, Michael C. Yip et al.· 2026 IEEE/ASME International...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.