ReCAST: Reward Credit Assignment across Timesteps for Online Diffusion Reinforcement
Improvements are shown that ReCAST yields improvements that generalize beyond the training rewards and support its core principle: assigning each reward greater weight at the denoising timesteps where its feedback is most informative.
Yi-Hang Chen, Yuan-Hao Ban, Kuei-Chun Kao et al.
· 0 citations