Skip to content
Review Open access

Reinforcement Learning for Diffusion Policies in Robotics: A Survey and State-Based Locomotion Reproduction

Aug 2026 · Robotics · 0 citations · 33 references

Abstract

Diffusion policies model multimodal robot action sequences, but behavioral cloning does not directly optimize task return. We present a structured scoping review of reinforcement learning for generative robot policies and a bounded state-based locomotion reproduction. Four documented routes yielded 178 records, 162 unique candidates, and an 84-study evidence map. Hierarchical rules distinguish 41 direct reward-driven studies from 32 adjacent robotic, eight alternative-generator, and three non-robotic studies; a five-axis taxonomy codes initialization/data, interaction regime, optimized object, credit assignment, and generator. Under a fixed-final evaluation protocol on the Datasets for Deep Data-Driven Reinforcement Learning (D4RL) 1.1 Hopper benchmark, five diffusion policy policy optimization (DPPO) fine-tuning seeds improved over their run-recorded behavior-cloning initializations by a mean of 1261.2 return, with a seed-level standard deviation of 125.5 and a 95% confidence interval of 1105.3–1417.1; the five runs link to two recorded behavior-cloning checkpoints. A Gaussian-policy control also improved after proximal policy optimization, so the gain was not diffusion-specific. A full-chain backpropagation adaptation exhibited clear seed-dependent variation, a matched action-divergence intervention did not establish causal critical timesteps, and reducing denoiser evaluations from 20 to 2 lowered A100 latency from 30.97 to 3.85 ms while substantially reducing normalized score. The experiments are limited to state-based locomotion and do not validate visual manipulation.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.