Skip to content

Binary Networks and Continual Learning for Pose Estimation from a Single Aerial Image

Aug 2026 · Unmanned Systems · pp. 1-10 · 0 citations

TL;DR

This work proposes a methodology using a binary network with a Continual Learning (CL) strategy to create an estimation model to create an estimation model during the same flight mission for Pose estimation using aerial images captured by UAVs.

Abstract

Pose estimation using aerial images captured by Unmanned Aerial Vehicles (UAVs) allows the localisation in GPS-denied scenarios. Several methods based on deep learning approaches with convolutional neural networks (CNN) have become tools for estimating localisation from images. However, building a model that can estimate the pose from a single image needs a large dataset and training time to obtain a result. Besides, the model can be inappropriate in assessing the correct pose in dynamic scenarios with multiple changes. Therefore, we propose a methodology using a binary network with a Continual Learning (CL) strategy to create an estimation model during the same flight mission. Also, we use a submap scheme and multiple models to acquire the UAV’s localisation into different parts of the trajectory. Finally, we use PoseNet, ORB-SLAM2 and single-model for comparison purposes in four scenarios, achieving a percentage error of 14% of the total trajectory and a processing time of 51 ms with our proposed approach.

View source

Similar papers

Preprint Aug 2026

Evaluation of Image Matching Methods for Visual Odometry on UAVs

Unmanned aerial vehicles (UAVs) are becoming a powerful tool for many environmental monitoring and transport applications. Yet, their reliance on Global Navigation Satellite System (GNSS) technology for navigation makes them susceptible to catastrophic failures in scenarios where the positioning signal is unavailable or disrupted. This work explores Visual Odometry (VO) as a crucial navigation component. Recently, numerous deep-learning-based methods for image matching have been proposed that are yet to be implemented in a fully-fledged VO system. In this paper, we evaluate recent state-of-the-art image matching methods for the task of VO for UAV position tracking, with a downwards-facing camera, on our synthetic dataset, and find that while the best results are generated by the recent RoMa matcher, SIFT features can outperform some recent state-of-the-art.

Gašper Spagnolo, L. Č. Zajc, Matej Dobrevski · 1 citation
Aug 2026

3D CNN and Study of Attention Mechanisms to Relative Camera Position

Calculating the relative position between images in Unmanned Aerial Vehicles (UAVs) is a core component in tasks such as visual odometry, SLAM, and state estimation. It enables the UAV to estimate its movement between frames using onboard sensors. In this paper, we present a spatio-temporal regression architecture combining 3D CNNs and attention mechanisms. The model takes a sequence of image frames and corresponding Inertial Measurement Unit (IMU) data for each frame as input and outputs the estimated relative position in meters between the images. The inclusion of IMU measurements is critical for improving robustness to motion blur, low-texture environments, and rapid maneuvers, as it provides complementary information about the UAV’s linear acceleration and angular velocity. The CNN layers extract compact spatio-temporal features from the video input, while the multihead attention layer captures temporal dependencies and contextual relations across time. The final Multi-Layer Perceptron (MLP) regresses these fused representations into a relative position estimate. We systematically evaluate different attention mechanisms: Masked, Causal, Local-window, and Linformer, to analyze their impact on accuracy and efficiency in relative position estimation. We demonstrate the effectiveness of our method on the TII Drone Racing and UZH-FPV datasets.

Nilda G. Xolo-Tlapanco, J. Martínez-Carranza · 0 citations
#machine learning Preprint Sep 2026

Infrastructure-based Monocular 3D Vehicle Localization Framework with Experimental Validation

This paper presents a one-stage learning framework that maps monocular roadside-camera images directly to vehicle states in a ground-fixed coordinate frame. Unlike conventional approaches that first detect vehicles in the image plane and subsequently apply geometric post-processing, the proposed method leverages features from a pretrained object detector to jointly estimate each vehicle's ground-plane position, dimensions, and yaw angle. The framework therefore uses visual features not only for vehicle detection but also for direct spatial and orientation estimation. To support model training and evaluation, we develop a data-collection and label-generation pipeline based on synchronized video from a roadside camera and an unmanned aerial vehicle (UAV). Acting as a temporary top-view sensing platform, the UAV provides vehicle trajectories, dimensions, and orientations, which are transformed into the ground-fixed coordinate frame and temporally aligned with the roadside-camera images to generate ground-truth labels. The framework is evaluated using data collected during multiple experiments at the Mcity Test Facility. Results show that the proposed method can recover vehicle trajectories and orientations from monocular roadside imagery without a separate geometric post-processing stage, demonstrating its potential as a scalable approach to infrastructure-based perception at urban intersections.

Akos T. Kopeczi-Bocz, Tian Mi, Gábor Orosz et al. · 0 citations
Preprint Aug 2026

AirAlign: Geometry-Aware Relative Pose Alignment for UAV Last-Meter Navigation

AirAlign is proposed, a framework for RGB-only image-pair relative pose alignment for UAVs, using a pretrained visual geometry reconstruction model as the backbone to extract geometry-aware features from source-target image pairs.

Jin-Yi Zhou, Shuo Feng, Yufei Wu et al. · 0 citations
Open access Jul 2026

Real-Time Temporally Consistent Monocular 6D UAV Pose Estimation for Onboard Aerial Perception

AeroMotion6D is proposed, a temporal transformer-based framework for monocular UAV 6D pose estimation from RGB video that consists of an adaptive context fusion mechanism that can incorporate past context information into the current estimation process and a persistent pose memory module that can convey pose-related information in two consecutive frames.

Mohammad Al Qaderi, M. Hayajneh, Alaa Alghazo et al. · 0 citations
Open access Aug 2026

Visual Autonomous Docking for Unmanned Surface Vehicles Using Lightweight Supervised Learning Framework

Autonomous docking is a core capability enabling full autonomy of unmanned surface vehicles (USVs), whose practical deployment demands visual pose estimation with high efficiency, temporal stability, and closed-loop control compatibility. This paper proposes a lightweight monocular visual docking perception framework based on MobileNetV2 and a temporal convolutional network (TCN). In this framework, a MobileNetV2 backbone is adopted to perform end-to-end regression of the USV’s relative pose with respect to the dock from a single monocular image, while a feature-level TCN module fuses sequential visual features across consecutive frames to enhance the short-term stability of pose estimation. To validate the performance and reliability of the proposed method, a high-fidelity simulation environment is established to conduct closed-loop USV docking tests. Comparative results demonstrate that the MobileNetV2 backbone reduces inference latency compared with the VGG19 architecture, and the embedded TCN module effectively suppresses inter-frame pose fluctuations and abnormal estimation jumps. The proposed method provides an efficient and temporally consistent visual perception solution for simulation-validated USV autonomous docking systems.

Jun-Yan He, Wei Liu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.