Extensive Exploration in Highway Overtaking Scenarios Using Hierarchical Reinforcement Learning
Abstract
Existing studies on deep reinforcement learning based driving controllers often focus on traffic scenarios with relatively simple patterns. This limits their ability to handle challenging highway overtaking scenarios with delayed long-term rewards and reduces the generalizability of the learned policy. This paper presents a hierarchical reinforcement learning framework for autonomous highway driving that decomposes delayed-reward highway overtaking decision making into interpretable subtasks. The framework contains a high-level controller for long-term planning and exploration and a low-level controller for detailed longitudinal and lateral control. To better expose the high-level controller to delayed rewards, the two controllers are trained separately in a two-step process. In the first step, the high-level controller is trained with a critic-gated goal completion mechanism and a fixed rule-based low-level motion planner. In the second step, the trained high-level policy guides the learning of the low-level controller. To evaluate long-horizon exploration, we design a speed-biased highway reward and a highway overtaking trap scenario involving both slower-moving vehicles and general traffic vehicles. We conduct the experiments in the highway-env simulation environment. Compared with Double DQN, the best-performing single-level controller design, our hierarchical framework improves trap-escape success from 0% to 100%, increases average speed from 10.63 m/s to 14.26 m/s, and increases the traveled distance from 248.17 m to 332.83 m during the fixed 50 decision steps long experimental episode window. Our method also achieves more reliable trap-escape performance than other hierarchical structures, including h-DQN and HIRO.