Skip to content

VINE: Taming Generative Control Policies for Reinforcement Learning

Jul 2026 · arXiv.org · Vol abs/2607.10369 · 0 citations · 66 references
Computer Science

TL;DR

VINE is proposed, an RL-oriented sampling method that enables stable end-to-end value-gradient optimization for flow-matching policies and achieves stable policy improvement and consistently outperforms state-of-the-art RL methods on the OGBench offline RL benchmark and real-world robotic manipulation task.

Abstract

Flow-matching policies have emerged as an effective policy parameterization for robot learning. They iteratively generate actions from noise, enabling highly expressive modeling of complex and multimodal action distributions. However, prior works observed that scaling these policies with value-gradient reinforcement learning (RL) often leads to training instability. Existing methods attribute this instability to iterative generation and therefore avoid end-to-end value-gradient optimization by sacrificing iterative generation, high expressiveness, or value-gradient optimization. Contrary to prior belief, we show the instability does not stem from iterative generation itself, but from the vanilla sampling strategy originally designed for behavior cloning, which becomes brittle under value-gradient RL. Motivated by this insight, we propose VINE, an RL-oriented sampling method that enables stable end-to-end value-gradient optimization for flow-matching policies. Instead of following a single flow trajectory, VINE reconstructs a new interpolation state at every denoising step, creating a stable differentiable path for value-gradient propagation while remaining compatible with the original flow-matching denoising process. As a result, VINE preserves the expressiveness and iterative generation of flow-matching without sacrificing end-to-end value-gradient optimization. Despite performing end-to-end backpropagation through all ten denoising steps, VINE achieves stable policy improvement and consistently outperforms state-of-the-art RL methods on the OGBench offline RL benchmark and real-world robotic manipulation task. Videos are available on our website: https://agibottech.github.io/vine.

View source

Similar papers

Preprint Aug 2026

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning

EvoHIL is presented, a unified framework that adapts the reward model, action generator, and visual do main within a staged human-in-the-loop learning process to improve task success, agreement with human-confirmation labels, motion smoothness, and completion time relative to human-in-the-loop and imitation baselines.

Shuoqing Zhang, Tongtong Cheng, Xiru Gao et al. · 0 citations
#machine learning Preprint Sep 2026

Multi-step Proximal Policy Improvement in Offline Reinforcement Learning

Offline reinforcement learning (RL) must reconcile two competing requirements: policy updates should stay near dataset-supported actions to keep value estimates reliable, yet meaningful gains often require moving beyond the behavior distribution. We develop a geometric view of offline actor updates by modeling policies as a probability manifold endowed with a chosen metric geometry. Under this lens, a broad class of offline actor objectives can be interpreted as a single proximal policy improvement step (SPI), i.e., an implicit discretization of a manifold gradient flow induced by a critic-defined energy. Building on this insight, we propose multi-step proximal policy improvement (MPI), a plug-in refinement mechanism that composes sequential re-centered proximal steps. MPI enables controlled policy improvement beyond dataset support while retaining proximal control at each refinement. The framework accommodates multiple policy geometries and admits practical instantiations for deterministic and diagonal-Gaussian policies. Experiments on D4RL benchmarks show that small numbers of MPI refinements improve strong offline baselines, including TD3+BC, ReBRAC, and IQL, on many tasks. Focused diagnostics further distinguish re-centered refinement from fixed-objective update scheduling and characterize limitations under critic error.

Soohyun Choi, Seonvin Cho, Songnam Hong · 0 citations
Aug 2026

Latent MeanFlow Policy Optimization for Offline Reinforcement Learning.

Offline reinforcement learning (RL) aims to derive effective policies from fixed datasets without environment interaction. While generative models such as diffusion and flow matching improve policy expressiveness, they often suffer from high computational cost due to multi-step sampling and limited representational richness from uninformative Gaussian priors. To address these challenges, we propose Latent MeanFlow Policy Optimization (LaMPO), a generative policy framework that formulates offline RL as latent generative policy optimization. LaMPO learns a behavior-conditioned latent distribution, which provides an informative prior and reduces the modeling burden of the generative process. Conditioned on this prior, the MeanFlow policy is trained to predict the average action velocity, enabling high-fidelity one-step action generation that avoids the cost of iterative sampling. Guided by the MeanFlow policy, the target policy is then optimized under a MeanFlow-induced behavior constraint, achieving effective policy improvement. We further establish a theoretical performance lower bound of the target policy relative to the MeanFlow policy. Extensive evaluations across 71 tasks in OGBench, D4RL, and real-world humanoid manipulation demonstrate that LaMPO achieves superior performance with high efficiency, yielding a 19% average improvement over existing state-of-the-art methods and an 81% success rate in real-world robotic tasks.

Tenglong Liu, Xin Xu, Yixing Lan et al. · 0 citations
Preprint Aug 2026

CoDrift: Compositional Drifting for Offline Reinforcement Learning

This work proposes CoDrift, a compositional framework for one-step generative policy learning that combines three objective-level fields into a unified policy field that compares favorably with state-of-the-art methods and achieves the best average rank in both settings.

Xiewei Ni, Ruo-Syuan Mei, Xiangyu Xu · 0 citations
Open access 2026

GCR-RL: Gradient Control Reward Shaping for Reinforcement Learning

This work introduces Gradient Control Rewards (GCR), an interpretable, control-inspired reward-design methodology for accelerating agent training by modulating the reward signal based on the temporal dynamics of system error, Inspired by classical control theory.

Anas Aburaya, H. Selamat, M. Muslim et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.