Gradient-free Robot Action generation via Combined diffusion-MPPI posterior mean Estimation (GRACE), which guides a pretrained diffusion policy with Model Predictive Path Integral (MPPI) control using only forward cost evaluations.
Abstract
Diffusion policies generate multimodal robot action sequences from demonstrations, but steering them toward deployment-time constraints typically relies on differentiable guidance costs. This excludes many practical safety constraints, such as binary collision checks, joint limits, and black-box rollout costs that are nondifferentiable. We propose Gradient-free Robot Action generation via Combined diffusion-MPPI posterior mean Estimation (GRACE), which guides a pretrained diffusion policy with Model Predictive Path Integral (MPPI) control using only forward cost evaluations. Building on the common score-ascent structure of diffusion and MPPI, GRACE constructs a cost-conditioned guidance posterior at each reverse step and estimates its mean with a single MPPI update centered at the diffusion reverse mean. For differentiable costs, GRACE recovers conventional gradient guidance under a first-order, matched-covariance approximation. GRACE attains higher success rates than diffusion-based and sampling-based baselines in simulation. On a real 7-DoF manipulator, GRACE avoids a deployment-time obstacle that the unguided prior collides with in every trial. Code and experiment videos are available at https://anonymous.4open.science/w/grace-70BB/.
Diffusion policies have demonstrated excellent performance in robotic control tasks, yet their reliance on 50 to 100 denoising steps impedes real-time deployment. Existing acceleration methods based on sampler improvements lack responsiveness to perceptual quality and cannot adaptively adjust computation. Moreover, sta...
Qi Chen, Xinyang Ren, Jiajun Xing et al.· IEEE Robotics and Automation...· 0 citations
Pretraining navigation diffusion policies rely on large-scale expert demonstrations. These data are typically generated by a fully-informed oracle planner suited to a single nominal robot. This limits the policy's generalization to diverse embodiments and challenging scenarios (e.g., escaping dead ends or detouring lon...
Tian-Yu Yang, Yiming Zeng, Wenzhe Cai et al.· arXiv.org· 0 citations
Temporal Policy is introduced, a generative framework based on stochastic interpolants that formulates action generation as a temporally coupled transport problem and bypasses the computational bottleneck of independent Gaussian priors, helping enable high-frequency, closed-loop control.
Dylan Miller, Martin Jägersand· arXiv.org· 0 citations
Decentralized multi-robot navigation is difficult when robots must act from local observations without centralized coordination or explicit inter-robot communication. A belief-driven hybrid reinforcement learning framework is evaluated for planar multi-robot navigation under partial observability. Each robot builds a c...
V. Malathi, Pramod Sreedharan, Rthuraj Puthiyaveedu Rajesh et al.· Robotics· 0 citations
Robot crowd navigation requires safe and efficient decision-making under dense, dynamic, and multimodal human--robot interactions. Existing reinforcement-learning methods typically output a single reactive action at each timestep, which limits their ability to represent diverse short-term avoidance strategies. We propo...
Diffusion policies model multimodal robot action sequences, but behavioral cloning does not directly optimize task return. We present a structured scoping review of reinforcement learning for generative robot policies and a bounded state-based locomotion reproduction. Four documented routes yielded 178 records, 162 uni...
Shihan Sun, Yinlong Liu· Robotics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.