Skip to content

Demonstration-Free Success-Probability Reward Learning for Generalist Robot Policies

Sep 2026 · 0 citations · 60 references
Computer Science

TL;DR

This work introduces eVTA, which learns success probabilities from mixed-quality policy rollouts through temporal-difference-style bootstrapping, without expert demonstrations or intermediate annotations, and introduces RL with Evolving Rewards (RLER), a closed-loop framework that adapts eVTA using newly collected rollouts as the policy evolves.

Abstract

Reinforcement learning (RL) enables generalist robot policies to improve through trial-and-error interaction, yet its effectiveness is fundamentally constrained by sparse task rewards. Existing general-purpose reward models typically alleviate this issue by learning task progress from expert demonstrations, but introduce a distribution mismatch with the mixed-quality rollouts encountered during policy optimization, making their estimates unreliable on suboptimal and failed behaviors from which the policy must learn. In this work, we introduce a demonstration-free reward learning paradigm where dense reward feedback can be learned directly from sparse task outcomes and policy experience. We theoretically show that terminal task outcomes implicitly define dense success-probability feedback at intermediate timesteps, which can be recursively learned through bootstrapping. Based on this insight, we introduce eVTA$_0$, which learns success probabilities from mixed-quality policy rollouts through temporal-difference-style bootstrapping, without expert demonstrations or intermediate annotations. We further introduce RL with Evolving Rewards (RLER), a closed-loop framework that adapts eVTA$_0$ using newly collected rollouts as the policy evolves. Experiments show that eVTA$_0$ provides more informative rewards than state-of-the-art reward models and achieves the best average policy performance across all LIBERO task suites under the same RL training budget, improving success rates by 5.4%-13.8% over the initial policy. In real-world manipulation, RLER further improves overall success rates by 20%-26%, with 35%-36% gains under out-of-distribution conditions. These results demonstrate the effectiveness of demonstration-free reward learning and adapting rewards as the policy evolves. Project webpage: https://duowuyms.github.io/evta0.

View source

Similar papers

Preprint Sep 2026

RoboFFT: Finetuning generative robot policy via online reinforcement learning with forward process

Generative models, such as diffusion and flow-based models, have shown strong promise for robot policy learning by capturing complex and multimodal action distributions from demonstrations. However, policies trained solely with imitation learning often suffer from imperfect demonstrations and distributional shifts, whi...

Yu Li, Sheng-He Hu, Yu-Han Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

From Imitation to Reward Discovery: On-Policy Warmup for Agentic RL

Reinforcement learning with a verifiable reward (RLVR) offers a scalable approach to training language-model agents, yet sparse outcome rewards can leave early training with little signal for policy improvement. We identify an On-Policy Acceleration Phenomenon: in our main comparisons, RLVR initialized with on-policy d...

Yi-Tong Qiao, Tian-Tian He, Lei Liu et al. · 0 citations
#machine learning Preprint Sep 2026

Reward as Observation: Learning Reward-Based Policies for Rapid Adaptation

This paper explores a reward-based policy to achieve zero-shot transfer between source and target environments with completely different observation spaces. While humans can demonstrate impressive adaptation capabilities, deep neural network policies often struggle to adapt to a new environment and require a considerab...

Morgan Byrd, Maks Sorokin, Robert Wright et al. · 0 citations
Preprint Sep 2026

Learning Robot Policies from Sparse Success Signals via STL-Guided Stein Variational Policy Gradient

Learning robot policies for tasks with sparse success signals is challenging when completion depends on coordinated actions, precise contact outcomes, or satisfying several conditions together. Intricate physical interactions with the world further complicate these requirements. Prior work using conventional reward sha...

Hong-Rui Zheng, C. Vasile, Antonio Loquercio et al. · 0 citations
#machine learning Preprint Sep 2026

ARS: Agentic Reward System for Robot Learning

Progress reward modeling is the problem of estimating how a robot's behavior changes task progress over time. Reliable estimation requires distinguishing meaningful state changes from failed attempts and task-irrelevant actions. We introduce the Agentic Reward System (ARS), an inference framework for progress reward mo...

Sheng-Miao Hu, Wei-Yi Lu, Ling-Bing Zeng et al. · 0 citations
Open access Sep 2026

Temporal Expert-Based Reward Learning for Inverse Reinforcement Learning

Learning reward functions from expert demonstrations removes the need for manual reward engineering in reinforcement learning applied to robotic manipulation. Existing methods, however, require trajectory quality annotations, episode success labels, or produce implicit rewards that are difficult to inspect. This paper...

Francisco José Naranjo Campos, Juan G. Victores, Almudena Alcaide et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.