Dreamer-SAC: Off-Policy Learning in Latent World Models for Sample-Efficient Autonomous Driving
The Dreamer-SAC framework, which integrates a recurrent state-space world model with an off-policy soft actor-critic algorithm trained directly in latent space, uses a combination of real interactions and short-horizon generated trajectories with n-step target estimation and multi-objective supervision.