From Segmentation to Navigation: Evaluating U-Net and SegNet with ResNet50, ResNeSt50d, and MobileNetV3 for Visual Semantic UAV Navigation
Abstract
Autonomous navigation remains one of the central challenges in deploying unmanned aerial vehicles (UAVs) for applications such as infrastructure inspection, last-mile delivery, and search-and-rescue, particularly in environments where reliable satellite positioning is degraded or unavailable. Visual semantic segmentation offers a practical foundation for such navigation by enabling a UAV to identify traversable regions, such as roads, directly from onboard camera imagery. This paper presents a comparative evaluation of two widely used encoder-decoder segmentation architectures, U-Net and SegNet, each paired with three backbone networks, ResNet50, ResNeSt50d, and MobileNetV3, for the task of binary road segmentation from aerial imagery. A custom dataset of approximately 2,500 image-mask pairs at 256×256 resolution was assembled by combining the UAVid dataset, synthetic imagery from GTAV, and manually annotated frames extracted from drone video footage. Each of the six architecture-backbone combinations was independently hyperparameter-tuned and trained for 50 epochs, then evaluated using Intersection over Union (IoU), Dice coefficient, precision, recall, and inference time. The resulting segmentation masks were further processed through a navigation pipeline consisting of morphological cleaning and filtering, skeletonization, and node detection, converting pixel-level predictions into a navigable path representation suitable for guiding a UAV along a road. These findings provide practical guidance for selecting segmentation models suited to real-time, vision-based UAV road-following navigation.