RefineFly: Failure-Aware Post-Training for Aerial Vision-Language Navigation
Abstract
Unmanned aerial vehicle vision-language navigation (UAV-VLN) requires agents to translate visual observations and language instructions into reliable flight actions in complex environments. Although recent end-to-end UAV vision-language-action (UAV-VLA) policies reduce reliance on separately designed perception, planning, and control modules, their behavior-cloning objectives provide limited supervision for errors arising during closed-loop execution. Reinforcement learning (RL) offers a promising solution, while limited interaction budgets, long-tailed scene distributions, and policy drift complicate sustained learning. To this end, we propose RefineFly, a failure-aware RL post-training framework for end-to-end UAV-VLA policies. Built on token-level proximal policy optimization (PPO), RefineFly maintains dynamic failure memory within a two-stage scene curriculum to sustain learning from unresolved tasks as training shifts from empirical to balanced scene sampling. Throughout this process, stage-specific reference regularization constrains policy deviation. Experiments on the TravelUAV benchmark demonstrate that RefineFly outperforms all comparison methods across all three splits, improving success rate over AerialVLA by 3.12 to 8.37 percentage points with a total rollout budget of about 30\% of the training-set size. Moreover, ablations reveal that the benefits of sustained failure learning vary across evaluation splits, while the two-stage scene curriculum improves overall performance, highlighting the importance of training distributions when learning from failures.