Calculating the relative position between images in Unmanned Aerial Vehicles (UAVs) is a core component in tasks such as visual odometry, SLAM, and state estimation. It enables the UAV to estimate its movement between frames using onboard sensors. In this paper, we present a spatio-temporal regression architecture combining 3D CNNs and attention mechanisms. The model takes a sequence of image frames and corresponding Inertial Measurement Unit (IMU) data for each frame as input and outputs the estimated relative position in meters between the images. The inclusion of IMU measurements is critical for improving robustness to motion blur, low-texture environments, and rapid maneuvers, as it provides complementary information about the UAV’s linear acceleration and angular velocity. The CNN layers extract compact spatio-temporal features from the video input, while the multihead attention layer captures temporal dependencies and contextual relations across time. The final Multi-Layer Perceptron (MLP) regresses these fused representations into a relative position estimate. We systematically evaluate different attention mechanisms: Masked, Causal, Local-window, and Linformer, to analyze their impact on accuracy and efficiency in relative position estimation. We demonstrate the effectiveness of our method on the TII Drone Racing and UZH-FPV datasets.
Nilda G. Xolo-Tlapanco, J. Martínez-Carranza· Unmanned Systems· 0 citations
We propose a novel methodology for autonomous landing zone detection in Micro Aerial Vehicles (MAVs) based on Vision Transformers (ViTs). The core contribution of this work lies in demonstrating that transformer-based architectures, originally developed for large-scale vision tasks, can be effectively adapted to safety-critical aerial robotics applications with limited training data. Unlike traditional Convolutional Neural Networks (CNNs), ViTs leverage self-attention mechanisms to model long-range spatial dependencies, enabling a more holistic understanding of scene geometry and surface suitability for landing. We systematically evaluate the proposed approach on aerial RGB images from a public dataset as well as on noisy depth images captured onboard a drone using a lightweight depth camera. Our results show that the ViT-based model consistently outperforms widely used CNN architectures, including ResNet, particularly in low-data regimes where generalization is crucial. Notably, the transformer model maintains strong robustness even when operating on degraded depth inputs. In addition to accuracy improvements, the proposed system achieves real-time performance, with an average inference time of [Formula: see text] ms on legacy GPU hardware. These findings highlight the practical feasibility and effectiveness of Vision Transformers for reliable, efficient MAV landing zone detection.
Victoria Eugenia Vazquez-Meza, J. Martínez-Carranza· Unmanned Systems· 0 citations
This work proposes a methodology using a binary network with a Continual Learning (CL) strategy to create an estimation model to create an estimation model during the same flight mission for Pose estimation using aerial images captured by UAVs.
A. Cabrera-Ponce, L. Rojas-Perez, Manuel Martin-Ortiz et al.· Unmanned Systems· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.