P3: Probabilistic Policy Propagation for Stable VAE-Based Robot Learning
This work introduces P^3 (Probabilistic Policy Propagation), a distribution-aware optimization framework for VAE-based policies that couples moment-based probabilistic method for stable and efficient learning with sampling-based calibration for robust policy behavior under latent uncertainty.