Skip to content

Author

J. Martínez-Carranza

We have 3 of 140 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Aug 2026

3D CNN and Study of Attention Mechanisms to Relative Camera Position

Calculating the relative position between images in Unmanned Aerial Vehicles (UAVs) is a core component in tasks such as visual odometry, SLAM, and state estimation. It enables the UAV to estimate its movement between frames using onboard sensors. In this paper, we present a spatio-temporal regression architecture combining 3D CNNs and attention mechanisms. The model takes a sequence of image frames and corresponding Inertial Measurement Unit (IMU) data for each frame as input and outputs the estimated relative position in meters between the images. The inclusion of IMU measurements is critical for improving robustness to motion blur, low-texture environments, and rapid maneuvers, as it provides complementary information about the UAV’s linear acceleration and angular velocity. The CNN layers extract compact spatio-temporal features from the video input, while the multihead attention layer captures temporal dependencies and contextual relations across time. The final Multi-Layer Perceptron (MLP) regresses these fused representations into a relative position estimate. We systematically evaluate different attention mechanisms: Masked, Causal, Local-window, and Linformer, to analyze their impact on accuracy and efficiency in relative position estimation. We demonstrate the effectiveness of our method on the TII Drone Racing and UZH-FPV datasets.

Nilda G. Xolo-Tlapanco, J. Martínez-Carranza · 0 citations
Aug 2026

Vision Transformer-Based Autonomous Safe Landing Zone Detection for MAVs

We propose a novel methodology for autonomous landing zone detection in Micro Aerial Vehicles (MAVs) based on Vision Transformers (ViTs). The core contribution of this work lies in demonstrating that transformer-based architectures, originally developed for large-scale vision tasks, can be effectively adapted to safety-critical aerial robotics applications with limited training data. Unlike traditional Convolutional Neural Networks (CNNs), ViTs leverage self-attention mechanisms to model long-range spatial dependencies, enabling a more holistic understanding of scene geometry and surface suitability for landing. We systematically evaluate the proposed approach on aerial RGB images from a public dataset as well as on noisy depth images captured onboard a drone using a lightweight depth camera. Our results show that the ViT-based model consistently outperforms widely used CNN architectures, including ResNet, particularly in low-data regimes where generalization is crucial. Notably, the transformer model maintains strong robustness even when operating on degraded depth inputs. In addition to accuracy improvements, the proposed system achieves real-time performance, with an average inference time of [Formula: see text] ms on legacy GPU hardware. These findings highlight the practical feasibility and effectiveness of Vision Transformers for reliable, efficient MAV landing zone detection.

Victoria Eugenia Vazquez-Meza, J. Martínez-Carranza · 0 citations
Aug 2026

Binary Networks and Continual Learning for Pose Estimation from a Single Aerial Image

This work proposes a methodology using a binary network with a Continual Learning (CL) strategy to create an estimation model to create an estimation model during the same flight mission for Pose estimation using aerial images captured by UAVs.

A. Cabrera-Ponce, L. Rojas-Perez, Manuel Martin-Ortiz et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.