Skip to content
Preprint

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification

Jul 2026 · 0 citations · 28 references
Computer Science

Abstract

We present HBPI-UCRL, a model-based algorithm for hierarchical reinforcement learning (HRL) that learns high-level and low-level policies in parallel. HBPI-UCRL exploits the fact that a high-level transition corresponds to a multi-step transition at the low level. We introduce two conditions on the low-level dynamics that are sufficient to make parallel HRL learnable. When these conditions hold, we prove that HBPI-UCRL has a polynomial sample complexity in the problem parameters. In the sparse-reward, goal-directed setting, our sample complexity upper bound for HBPI-UCRL is strictly lower than that of its non-hierarchical counterpart, providing theoretical justification for the empirical success of HRL.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.