Skip to content

ReCAST: Reward Credit Assignment across Timesteps for Online Diffusion Reinforcement

Sep 2026 · 0 citations · 50 references
Computer Science

TL;DR

Improvements are shown that ReCAST yields improvements that generalize beyond the training rewards and support its core principle: assigning each reward greater weight at the denoising timesteps where its feedback is most informative.

Abstract

Training diffusion models with multiple rewards requires distinguishing user preference from reward informativeness. User preference determines how much each reward should contribute to the overall objective; reward informativeness determines when its feedback is useful during denoising. Some rewards can meaningfully evaluate a sample as soon as global structure emerges, but others become informative only when the sample is nearly clean. To address both questions jointly, we propose ReCAST (Reward Credit ASsignment across T}imesteps), the first method, to our knowledge, for per-reward, timestep-dependent credit assignment in diffusion reward fine-tuning. ReCAST separates user preferences from temporal allocation through a reward-by-timestep weight matrix $W$, whose row sums match the user-specified reward budgets $\lambda$, while its column sums are equal, assigning the same total weight to each denoising step. Under these marginal constraints, ReCAST allocates weight according to each reward's informativeness, quantified by its R\'enyi discriminability gain at each step. These gains telescope to the total discriminability between the reward-induced positive policy and the current policy, providing a basis for temporal credit assignment. We evaluate ReCAST by training SD3.5-Medium under two distinct four-reward settings, each across five reward budgets $\lambda$. ReCAST improves the training rewards in one setting and matches them in the other, improves every held-out judge in both, and is preferred by an independent LLM-as-a-Judge. Together, these results show that ReCAST yields improvements that generalize beyond the training rewards and support its core principle: assigning each reward greater weight at the denoising timesteps where its feedback is most informative.

View source

Similar papers

Preprint Aug 2026

Latent Reward Registers for Diffusion Preference Alignment

This work proposes Latent Reward Registers, a mechanism that estimates terminal preference directly from intermediate noisy latents, and achieves significant reward improvement with a favorable reward-quality balance against training-free baselines.

Yuan-Shen Guan, Zipeng Feng, Cheng-Ru Song et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Diffusion Reward Models

Reward models underpin the alignment of large language models, yet the dominant designs reduce each prompt--response pair to a point estimate or to a distribution from a fixed parametric family. This is at odds with human preference, which is inherently multimodal: the same response can be reasonably judged in many way...

Xiang-Yang Wang, Bing-Xiang He, Ze-Yuan Liu et al. · 0 citations
#machine learning Preprint Sep 2026

When Sparse Reward Meets Dense Distillation: Training Dynamics of On-Policy Distillation

The cross-signal NTK is introduced, a token-level statistic that measures the alignment between reward and distillation gradients at position n and an empirical threshold beyond which naive mixing can lead to persistent training collapse is revealed.

Xin-Ke Jiang, Tao Feng, Zhi-Bang Yang et al. · 0 citations
#machine learning Preprint Sep 2026

IncentRL: The Trade-Off Between Preference Guidance and Task Performance

Preference-based reward shaping can guide reinforcement learning, but adding preference signals to the reward may unintentionally change the task being optimized. We address this problem with IncentRL, a framework that introduces preference guidance while explicitly characterizing its effect on external-task performanc...

Xue-Ning Wu, Yan-Lan Kang, Shen Yin · 0 citations
Preprint Sep 2026

Tail-Likelihood Reinforcement Learning

Tail-Likelihood Reinforcement Learning (TailRL), which maximizes the log-probability of exceeding a randomly chosen reward threshold, which gives more weight to rare, high-reward rollouts and can be interpreted as a mixture of Best-of-k gradients.

Shrinivas Ramasubramanian, Daman Arora, Fahim Tajwar et al. · 1 citation

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.