DIA: Denoising Intermediate Advantage for Diffusion Policy Optimization
DIA is introduced, a policy-gradient method that learns a value function over partially denoised actions and uses it to construct a denoising level advantage for each step of the generative process, enabling it to discover more effective and efficient task-level strategies and subtask sequences that baseline methods fa...