Skip to content
Review Open access

Deep Learning on the Wing: A Survey of Resource-Efficient Object Detection and SLAM for Autonomous UAVs

2026 · International Conference on Data Technologies and Applications · pp. 243-248 · 0 citations · 15 references
Computer Science

TL;DR

A systematic classification of recent breakthroughs in real-time perception, specifically evaluating stereo image processing and Simultaneous Localization and Mapping through the lens of computational economy, and analyzes the efficacy of specialized optimization frameworks designed to maximize inference speed on embedded CPU/GPU architectures.

Abstract

: The evolution of Unmanned Aerial Vehicles (UAVs) into self-governing robotic entities is currently limited by the high computational overhead of modern neural networks. This review investigates the intersection of high-fidelity deep learning and the stringent resource limitations of edge-based aerial hardware. We present a systematic classification of recent breakthroughs in real-time perception, specifically evaluating stereo image processing and Simultaneous Localization and Mapping (SLAM) through the lens of computational economy. Drawing on extensive industrial experience in drone manufacturing and AI department leadership, this paper analyzes the efficacy of specialized optimization frameworks—such as multi-threaded frame tiling and hardware-concurrency mapping—designed to maximize inference speed on embedded CPU/GPU architectures. We further examine the role of spatio-temporal modeling and LSTM-based architectures in navigating unpredictable environments, while synthesizing the requirements for safety-critical, responsible AI deployment. By aligning theoretical algorithmic pruning with the practical realities of the product lifecycle, this survey provides a definitive technical roadmap for engineers and researchers aiming to achieve robust, on-board autonomy in the next generation of intelligent flight systems.

Read PDF

Similar papers

Preprint Aug 2026

DPNet: Efficient Dead-End Prediction and Avoidance for Vision-Based UAV Navigation

Vision-based Unmanned Aerial Vehicles (UAVs) often suffer from navigation failures in dead ends due to limited sensing accuracy and range. To address this challenge, this paper proposes a systematic solution for efficient dead-end prediction and avoidance. The proposed method introduces a lightweight neural network to predict the relative distance and bearing of potential dead ends within the current field of view using RGB-D inputs. These predictions prune a predefined, compact trajectory library, enabling the planner to proactively avoid dead ends while maintaining navigational smoothness. Notably, our approach transfers across real-world scenarios without manual annotation or fine-tuning on real-world data. The system achieves high-frequency replanning at 50 Hz onboard. Extensive simulation benchmarks demonstrate superior performance in success rate, flight time, and trajectory length, and real-world experiments further validate its effectiveness in complex scenarios.

Ruibin Zhang, Lun Pan, Zelong Xia et al. · 0 citations
Preprint Aug 2026

CoNav-UAV: Cooperative Dual-Altitude Aerial Navigation via Stackelberg Learning

Target-oriented vision-and-language navigation (VLN) on aerial platforms is attracting growing attention for missions such as disaster rescue, infrastructure inspection, and security patrol. In this task, an unmanned aerial vehicle (UAV) needs to locate targets given only a concise description of their appearance and surroundings. This requires global exploration and grounding as well as collision-free close-range approach, two interleaved processes difficult to reconcile within a single agent. Most existing methods transfer the ground VLN paradigm to a low-altitude UAV and compensate for its inefficient exploration with external assistance. A recent attempt deploys two UAVs at complementary altitudes yet still relies on privileged information and trains its two agents independently, precluding any mutual adaptation essential for cooperation. Here we propose CoNav-UAV, which explicitly models the task as a Stackelberg game between a high-altitude leader and a low-altitude follower, with the system operating on onboard visual and linguistic inputs alone. To solve this game, we introduce Iterative Stackelberg Learning. The leader's high-level vision-language reasoning is refined via memory-based in-context learning, while the follower's precise motion control is updated via DAgger-style expert distillation. The alternation drives both agents toward a Stackelberg equilibrium. CoNav-UAV consistently outperforms single- and dual-agent baselines across three high-fidelity urban scenes from the AerialVLN benchmark. Success rate improves by up to 30.8 points on the learning scene, and 9.0 points under cross-scene transfer while using about 3x less adaptation data. Further analyses validate the complementary gains of the leader and follower updates and reveal robust gains yet distinct learning dynamics across VLM backbones.

Junru Song, Wenhao Zhang, Yang Yang et al. · 0 citations
Preprint Aug 2026

AeroDPO: Unleashing Lightweight UAV Navigation with High-Fidelity Perception and Automated Preference Optimization

Vision-Language Navigation for Unmanned Aerial Vehicles (UAV-VLN) requires rapid and reactive control in complex 3D environments. Recent minimalist end-to-end paradigms show great promise but typically rely on massive language models containing billions of parameters, incurring prohibitive latency for real-world edge deployment. In this paper, we challenge this parameter-heavy reliance. Comprehensive cross-scale evaluations reveal the critical insight that perception quality fundamentally outweighs language reasoning capacity. We demonstrate that a lightweight 2B model equipped with high-fidelity visual inputs completely matches the overall success rates of massive 7B baselines. However, this minimalist policy exposes a fundamental robustness flaw inherent to pure Behavior Cloning (BC). Lacking explicit negative feedback, the agent fails to internalize robust spatial constraints and exhibits alarming collision rates in out-of-distribution (OOD) scenarios. To overcome this vulnerability without relying on unscalable human annotations, we propose AeroDPO, a zero-cost automated Direct Preference Optimization pipeline driven by deterministic physical simulation state rollback. Upon detecting collisions, the system autonomously rewinds the environment to extract causal reasoning errors as rejected actions, applies decoupled privileged interventions to synthesize collision-avoidance preferred maneuvers, and leverages an offline vision language inspector to filter visual ambiguities. By equipping our 2B model with this automated data flywheel, AeroDPO boosts success rates to 49.16% on unmapped scenarios while drastically suppressing collision rates, establishing a new SOTA for autonomous aerial agents.

Peng Xu, Chengcheng Wang, Shaohua Wan · 0 citations
Review Open access Jul 2026

Advances in Trajectory Prediction for High-Speed UAVs: A Review

Comparative analysis reveals that no single technical route can fully address the coupled challenges of uncertainty, accuracy, and real-time performance, underscoring that hybrid frameworks are essential for balancing these competing requirements.

Wenqin Han, Shuangxi Liu, Xianyu Wu et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.