VINE is proposed, an RL-oriented sampling method that enables stable end-to-end value-gradient optimization for flow-matching policies and achieves stable policy improvement and consistently outperforms state-of-the-art RL methods on the OGBench offline RL benchmark and real-world robotic manipulation task.
Abstract
Flow-matching policies have emerged as an effective policy parameterization for robot learning. They iteratively generate actions from noise, enabling highly expressive modeling of complex and multimodal action distributions. However, prior works observed that scaling these policies with value-gradient reinforcement learning (RL) often leads to training instability. Existing methods attribute this instability to iterative generation and therefore avoid end-to-end value-gradient optimization by sacrificing iterative generation, high expressiveness, or value-gradient optimization. Contrary to prior belief, we show the instability does not stem from iterative generation itself, but from the vanilla sampling strategy originally designed for behavior cloning, which becomes brittle under value-gradient RL. Motivated by this insight, we propose VINE, an RL-oriented sampling method that enables stable end-to-end value-gradient optimization for flow-matching policies. Instead of following a single flow trajectory, VINE reconstructs a new interpolation state at every denoising step, creating a stable differentiable path for value-gradient propagation while remaining compatible with the original flow-matching denoising process. As a result, VINE preserves the expressiveness and iterative generation of flow-matching without sacrificing end-to-end value-gradient optimization. Despite performing end-to-end backpropagation through all ten denoising steps, VINE achieves stable policy improvement and consistently outperforms state-of-the-art RL methods on the OGBench offline RL benchmark and real-world robotic manipulation task. Videos are available on our website: https://agibottech.github.io/vine.
EvoHIL is presented, a unified framework that adapts the reward model, action generator, and visual do main within a staged human-in-the-loop learning process to improve task success, agreement with human-confirmation labels, motion smoothness, and completion time relative to human-in-the-loop and imitation baselines.
Shuoqing Zhang, Tongtong Cheng, Xiru Gao et al.· 0 citations
Offline reinforcement learning (RL) must reconcile two competing requirements: policy updates should stay near dataset-supported actions to keep value estimates reliable, yet meaningful gains often require moving beyond the behavior distribution. We develop a geometric view of offline actor updates by modeling policies as a probability manifold endowed with a chosen metric geometry. Under this lens, a broad class of offline actor objectives can be interpreted as a single proximal policy improvement step (SPI), i.e., an implicit discretization of a manifold gradient flow induced by a critic-defined energy. Building on this insight, we propose multi-step proximal policy improvement (MPI), a plug-in refinement mechanism that composes sequential re-centered proximal steps. MPI enables controlled policy improvement beyond dataset support while retaining proximal control at each refinement. The framework accommodates multiple policy geometries and admits practical instantiations for deterministic and diagonal-Gaussian policies. Experiments on D4RL benchmarks show that small numbers of MPI refinements improve strong offline baselines, including TD3+BC, ReBRAC, and IQL, on many tasks. Focused diagnostics further distinguish re-centered refinement from fixed-objective update scheduling and characterize limitations under critic error.
Offline reinforcement learning (RL) aims to derive effective policies from fixed datasets without environment interaction. While generative models such as diffusion and flow matching improve policy expressiveness, they often suffer from high computational cost due to multi-step sampling and limited representational richness from uninformative Gaussian priors. To address these challenges, we propose Latent MeanFlow Policy Optimization (LaMPO), a generative policy framework that formulates offline RL as latent generative policy optimization. LaMPO learns a behavior-conditioned latent distribution, which provides an informative prior and reduces the modeling burden of the generative process. Conditioned on this prior, the MeanFlow policy is trained to predict the average action velocity, enabling high-fidelity one-step action generation that avoids the cost of iterative sampling. Guided by the MeanFlow policy, the target policy is then optimized under a MeanFlow-induced behavior constraint, achieving effective policy improvement. We further establish a theoretical performance lower bound of the target policy relative to the MeanFlow policy. Extensive evaluations across 71 tasks in OGBench, D4RL, and real-world humanoid manipulation demonstrate that LaMPO achieves superior performance with high efficiency, yielding a 19% average improvement over existing state-of-the-art methods and an 81% success rate in real-world robotic tasks.
Tenglong Liu, Xin Xu, Yixing Lan et al.· IEEE Transactions on Pattern...· 0 citations
This work proposes CoDrift, a compositional framework for one-step generative policy learning that combines three objective-level fields into a unified policy field that compares favorably with state-of-the-art methods and achieves the best average rank in both settings.
RoMAN-Flow (Robotic Manipulation with Autoregressive Normalizing Flows), an offline reinforcement learning framework that makes AR-NF policies practical for robotic manipulation by addressing this sampling bottleneck in both stages.
Shaoxuan Wang, Guangting Zheng, Rui Huang et al.· 0 citations
This work introduces Gradient Control Rewards (GCR), an interpretable, control-inspired reward-design methodology for accelerating agent training by modulating the reward signal based on the temporal dynamics of system error, Inspired by classical control theory.
Anas Aburaya, H. Selamat, M. Muslim et al.· IEEE Access· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.