EmbodiedMind: Adaptive Data Curation and Prefix-Tree Reinforcement Learning for Efficient Embodied Intelligence
Trie-GRPO, a novel reinforcement learning algorithm based on action prefix trees, which enables step-level advantage estimation, is introduced, which resolves the credit assignment problem by isolating intermediate correct decisions from downstream errors, while effectively balancing exploration efficiency and depth co...