Teacher-Guided Asymmetric Reinforcement Learning for End-to-End Visual Navigation of UAVs
Abstract
Autonomous navigation of low-altitude unmanned aerial vehicles (UAVs) in cluttered environments is challenging due to partial observability, limited onboard perception, and inefficient exploration in end-to-end reinforcement learning. This paper proposes a teacher-guided asymmetric reinforcement learning framework for end-to-end visual navigation of low-altitude UAVs in the Isaac Sim 5.1 environment. A privileged teacher policy is first trained using obstacle-state information to acquire reliable navigation priors. A deployable student policy is then learned with an asymmetric actor-critic architecture, where the actor takes depth images and proprioceptive states as input, while the critic uses privileged information during training. To improve policy transfer, an annealed knowledge distillation strategy is adopted: the student is strongly guided by the teacher in the early stage, and the guidance is gradually removed to enable autonomous reinforcement refinement. Experimental results show that the proposed method achieves higher success rates, lower collision rates, and faster convergence than baseline methods.