Skip to content

Bellman Diffusion Models for Offline Reinforcement Learning

· 0 citations · 20 references

TL;DR

This work explores using diffusion models as a representation for the state successor measure and finds that enforcing the Bellman flow constraints on a diffusion model leads to a temporal difference update on the predicted noise, similar to the standard TD-learning update on the predicted reward.

View source

Similar papers

#machine learning Preprint Oct 2026

Taylor Representations for Model-Free RL in Networked MDPs

In Networked Markov Decision Processes, transition dynamics are often unknown and the state--action space grows rapidly with the number of agents. In this setting, Taylor representations naturally approximate $Q$-functions, but a naive order-$n$ expansion over $N$ agents requires $\Theta(N^n)$ coefficients. We justify...

Salah Chikhi, Abdelhaq Chaoui, A. Ozdaglar et al. · 0 citations
Preprint Aug 2026

Parameter Exploration for RLVR via Variational Learning

Evidence that parameter-space exploration can improve reinforcement learning for LLMs is presented, and a family of methods called Perturbed Parameter Policy Optimization (3PO) is introduced which use different sampling strategies and different rollout grouping for reward estimation.

Vatsal Venkatkrishna, Nico Daheim, Iryna Gurevych · 0 citations
#machine learning Preprint Aug 2026

Designing Reinforcement Learning for Diffusion Models: A Unified Path-Space View

The empirical gap between these method families is identified as a variance-reduction effect rather than a difference in RL principle, and a multi-sample KDE value-gradient estimator that reuses rollout groups, together with scale-bounded weight families that retain stable existing recipes while excluding singular ones...

Yi-Xian Xu, Yuanrui Zhang, Shengjie Luo et al. · 2 citations
#artificial intelligence Review Oct 2026

Optimal Transport Meets Reinforcement Learning: A Survey

Reinforcement learning (RL) algorithms frequently compare probability distributions, such as state visitation distributions induced by policies and experts, action distributions from learned policies and offline datasets, or transition distributions from learned models and environments. However, commonly used divergenc...

Yu-Jie Zhu, Charles A. Hepburn, Matthew Thorpe et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Noisy-Space Policy Gradient for Diffusion Policies in Offline Reinforcement Learning

Diffusion policies offer a powerful and expressive parameterization for continuous control. Yet, their integration with reinforcement learning remains conceptually and algorithmically challenging. In this work, we address this gap by introducing a noisy-space action-value (Q-)function that assigns values to diffusion l...

Mahmoud Selim, Cristina Cipriani, K. H. Johansson · 0 citations
#machine learning Preprint Oct 2026

Reinforcement Learning for Hierarchical Reasoning Rewards: Minimax-Optimal Rates with Transformers

Reinforcement learning (RL) has become a standard tool for post-training language models on reasoning tasks, where the policy is updated by reward feedback while exploring the space of responses. Despite its empirical success, theoretical understanding of RL post-training remains limited, in particular of why on-policy...

Naoki Nishikawa, Taiji Suzuki · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.