Bellman Diffusion Models for Offline Reinforcement Learning
This work explores using diffusion models as a representation for the state successor measure and finds that enforcing the Bellman flow constraints on a diffusion model leads to a temporal difference update on the predicted noise, similar to the standard TD-learning update on the predicted reward.