2026· Annual Meeting of the Association for Computational Linguistics· pp. 39277-39307· 1 citation· 50 references
Computer Science
TL;DR
Experimental results show that ORBIT achieves SOTA performance on EB-ALFRED, outper-forming all closed-source and online-RL-based methods, while being substantially more effi-cient in training speed and computational cost, remaining robust to sub-optimal expert trajectories, and exhibiting strong generalization to unseen environments.
Abstract
Embodied planning requires agents to make coherent multi-step decisions based on dynamic visual observations and verbal goals. While recent vision-language models (VLMs) excel at static perception tasks, they struggle in interactive environments. Reinforcement learning (RL) offers a natural way to address this limitation, yet online RL approaches suffer from costly interaction and sparse rewards in embodied settings. This paper introduces ORBIT , an O n-policy R einforcement fine-tuning (RFT) framework with offline rewards for Em B od I ed T ask Planning, that preserves the generalization benefits of RFT while addressing the challenges of costly interaction and sparse rewards, supported by solid theoretical guarantees. Our approach is evaluated on EmbodiedBench, a recent benchmark for interactive embodied tasks, covering both in-domain and out-of-domain scenarios. Experimental results show that ORBIT achieves SOTA performance on EB-ALFRED, outper-forming all closed-source and online-RL-based methods, while being substantially more effi-cient in training speed and computational cost, remaining robust to sub-optimal expert trajectories, and exhibiting strong generalization to unseen environments. We released all code and data at https://github.com/mail-taii/Reinforced-Reasoning-for-Embodied-Planning
A coherent map of the rapidly expanding landscape of visual RL is provided to provide researchers and practitioners with a coherent map of the rapidly expanding landscape of visual RL and to highlight promising directions for future inquiry.
A unified framework that combines centralized training with decentralized execution (CTDE) and a Hybrid Reward Architecture (HRA) is introduced that enables multiple actors to share a centralized multi-head critic and substantially improves both sample efficiency and policy performance.
Changhao Li, Yifang Zhang, Heng Zhang et al.· 0 citations
EvoHIL is presented, a unified framework that adapts the reward model, action generator, and visual do main within a staged human-in-the-loop learning process to improve task success, agreement with human-confirmation labels, motion smoothness, and completion time relative to human-in-the-loop and imitation baselines.
Shuoqing Zhang, Tongtong Cheng, Xiru Gao et al.· 0 citations
ExToken is introduced, a simple yet general framework that condition VLA policies on discrete behavioral priors derived from offline demonstrations for structured exploration that consistently accelerates convergence, improves task performance, and exhibits strong robustness under highly constrained interaction budgets.
Yilun Kong, Yunpeng Qing, Guozheng Ma et al.· 0 citations
A data-driven scoping review of 130 studies published between 2020 and 2026, following PRISMA-ScR guidelines, to systematically map the landscape of long-horizon RL for robotic manipulation and presents a gap atlas that identifies underexplored research directions across methodological and experimental dimensions.
Matthew Acs, Xiangnan Zhong· Discover Robotics· 0 citations
This research bridges theoretical foundations of reinforcement learning and graph-based memory with autonomous agent workflows, and offers a practical, scalable reference framework for developing artificial intelligence technologies in complex, multi-step autonomous systems.