This work explores using diffusion models as a representation for the state successor measure and finds that enforcing the Bellman flow constraints on a diffusion model leads to a temporal difference update on the predicted noise, similar to the standard TD-learning update on the predicted reward.
In Networked Markov Decision Processes, transition dynamics are often unknown and the state--action space grows rapidly with the number of agents. In this setting, Taylor representations naturally approximate $Q$-functions, but a naive order-$n$ expansion over $N$ agents requires $\Theta(N^n)$ coefficients. We justify...
Salah Chikhi, Abdelhaq Chaoui, A. Ozdaglar et al.· 0 citations
Evidence that parameter-space exploration can improve reinforcement learning for LLMs is presented, and a family of methods called Perturbed Parameter Policy Optimization (3PO) is introduced which use different sampling strategies and different rollout grouping for reward estimation.
The empirical gap between these method families is identified as a variance-reduction effect rather than a difference in RL principle, and a multi-sample KDE value-gradient estimator that reuses rollout groups, together with scale-bounded weight families that retain stable existing recipes while excluding singular ones...
Yi-Xian Xu, Yuanrui Zhang, Shengjie Luo et al.· 2 citations
Reinforcement learning (RL) algorithms frequently compare probability distributions, such as state visitation distributions induced by policies and experts, action distributions from learned policies and offline datasets, or transition distributions from learned models and environments. However, commonly used divergenc...
Yu-Jie Zhu, Charles A. Hepburn, Matthew Thorpe et al.· 0 citations
Diffusion policies offer a powerful and expressive parameterization for continuous control. Yet, their integration with reinforcement learning remains conceptually and algorithmically challenging. In this work, we address this gap by introducing a noisy-space action-value (Q-)function that assigns values to diffusion l...
Mahmoud Selim, Cristina Cipriani, K. H. Johansson· 0 citations
Reinforcement learning (RL) has become a standard tool for post-training language models on reasoning tasks, where the policy is updated by reward feedback while exploring the space of responses. Despite its empirical success, theoretical understanding of RL post-training remains limited, in particular of why on-policy...
Naoki Nishikawa, Taiji Suzuki· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.