Skip to content

ActiveFly-Bench: Aligning Embodied Question Answering with Vision-Language-Action for Aerial Embodied Perception

Jul 2026 · arXiv.org · Vol abs/2607.10180 · 1 citation · 50 references
Computer Science

TL;DR

The first benchmark to bridge cyberspace reasoning and physical-world interaction for UAV embodied perception is introduced, and ActiveFly, a closed-loop UAV agent that integrates visual-language reasoning with fine-grained control, is developed and deployed on a physical UAV platform.

Abstract

We introduce ActiveFly-Bench, the first benchmark to bridge cyberspace reasoning and physical-world interaction for UAV embodied perception. The benchmark decomposes active perception into three hierarchical tasks: Aerial Embodied Question Answering (Air-EQA), Observation Behavior Planning (OBP), and Fine-grained Language-guided UAV Control (FLUC), explicitly connecting high-level task understanding, behavior planning, and low-level control. The datasets are collected from both real-world and simulated outdoor environments for training and evaluation. We further develop ActiveFly, a closed-loop UAV agent that integrates visual-language reasoning with fine-grained control, and deploy it on a physical UAV platform. Experiments with representative VLMs and VLA models show that current UAV agents still struggle with behavior planning, viewpoint adjustment, and robust task completion in active perception. These results establish ActiveFly-Bench as a new testbed for embodied aerial intelligence.

View source

Similar papers

Preprint Aug 2026

AeroGround: A Comprehensive Benchmark for Aerial-Ground Collaborative Reasoning

By systematically revealing the strengths and limitations of existing models in aerial-ground collaborative reasoning, AeroGround provides a foundation for developing more capable aerial-ground collaborative embodied intelligence systems.

Shenghong Yi, Lin Zhang, Muzian Li et al. · 0 citations
Review Open access Sep 2026

Large Language Models for UAV Autonomy from a Perception–Cognition–Action Perspective

Deployable autonomy remains a key challenge for unmanned aerial vehicles (UAVs) operating in open-ended missions. Large language models (LLMs) and their multimodal variants, which can process visual and other sensory inputs, have introduced new capabilities for semantic perception, task reasoning, and language-conditioned control. However, these capabilities do not by themselves produce flight-ready autonomy. We structure our analysis around a Perception–Cognition–Action (P–C–A) framework. At each layer, we identify the capabilities contributed by LLM-based components and examine how they connect to existing flight modules through input specifications, output representations, architectural coupling patterns, and safety mechanisms. Across the surveyed systems, LLMs extend UAV autonomy beyond fixed perception categories, scripted task plans, and pre-programmed controllers. However, field deployment depends on whether model outputs can be transformed into representations that downstream modules can parse, verify, and safely execute. Without adequate validation, captions, task plans, code, waypoints, and control commands may become failure points that propagate across the P–C–A loop. Our analysis highlights structured output contracts, independent safety barriers, and deterministic fallback mechanisms as key design elements for the reliable integration of LLM capabilities into UAV platforms.

Ting-Quan Xiong, Jianning Zhan, Qiu-Wei Deng et al. · 0 citations
Jul 2026

Zero-Shot Mission-Level Evaluation for Aerial MLLM Agents

This work introduces MissionBench, a benchmark for mission-level evaluation of MLLMs in aerial 3D environments, and shows that mission-level competence requires coordinating multiple capabilities beyond spatial perception, including multi-step planning and adaptive reasoning.

Suman Navaratnarajah, Taehyoung Kim, Jona Ruthardt et al. · 0 citations
#artificial intelligence Preprint Aug 2026

LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation

Embodied navigation requires agents to translate heterogeneous goals and visual observations into actions across tasks, environments, and robot embodiments. Modern vision-language models (VLMs) already encode spatial priors for visual grounding, spatial reasoning, and pointing, but these capabilities are rarely elicited directly for robot control. Existing navigation systems instead rely on task- or embodiment-specific components, fragmenting perception, reasoning, and action while offering limited generalization. Here we present LightNav-0, a compact generalist embodied navigation model that elicits the spatial intelligence of a pretrained VLM and aligns it with navigation, without task-specific prediction heads. LightNav-0 represents diverse navigation tasks through a unified token interface: dual-channel pointing expresses task-, scene-, and embodiment-agnostic spatial intent, while a residual vector-quantized action tokenizer maps this intent to precise, embodiment-specific trajectories. Together with temporally aware visual history compression, ER mid-training, supervised fine-tuning, and reinforcement learning, this formulation supports instruction following, open-vocabulary object navigation, and visual tracking within a single model. The navigation training corpus spans 2K+ scenes and 4K+ hours of embodied navigation data. LightNav-ER, the embodied-reasoning checkpoint used to initialize LightNav-0, attains the highest complete-set average across 8 embodied-reasoning benchmarks, while LightNav-0 achieves state-of-the-art monocular success rates across all 10 public navigation simulation settings. Real-world evaluations further demonstrate zero-shot generalization across robot embodiments, diverse scenes, and static and dynamic targets. These results establish compact VLMs as a unified and transferable backbone for generalist embodied navigation.

Shao-An Wang, Ao-Cheng Luo, Fei Huang et al. · 0 citations
Preprint Sep 2026

MobileVLA-R1 2.0: RL-Enhanced Reasoning for Mobile Robot Control

Grounding natural-language instructions into reliable and executable actions remains a fundamental challenge for vision-language-action (VLA) systems on mobile robots, due to the persistent gap between high-level semantic reasoning and low-level locomotion and manipulation control. Existing approaches often rely on implicit reasoning or monolithic action prediction, making it difficult to maintain coherent long-horizon decision making while producing precise and adaptable robot actions. To address this challenge, we propose MobileVLA-R1 2.0, an RL-enhanced VLA framework that explicitly couples structured embodied reasoning with executable mobile robot control. The framework learns multi-granularity reasoning over embodied trajectories through supervised Chain-of-Thought (CoT) alignment and reinforcement learning, improving reasoning-to-action consistency beyond purely behavioral supervision. To support both locomotion and manipulation, we further introduce a reasoning-conditioned action decoder that maps multimodal reasoning representations to task-level action targets, which are subsequently translated into embodiment-specific commands by robot controllers. This design provides a unified perception-reasoning-action interface while decoupling high-level action generation from robot-specific actuation. We conduct extensive evaluations on language-guided navigation, quadruped control, and humanoid mobile manipulation, covering VLN-CE, QUARD, and real-world deployments on Unitree Go2 and G1 robots. MobileVLA-R1 2.0 consistently outperforms strong VLA baselines, achieving an average 1.6 point improvement in SR on VLN-CE and a 10.0 point improvement in full-task success on real-world G1 mobile manipulation tasks over MobileVLA-R1, while demonstrating robust long-horizon instruction following and closed-loop execution across different robotic platforms.

Ting Huang, Yue Huang, Ze-Yu Zhang et al. · 0 citations
Preprint Aug 2026

Advancing MLLM-based UAV Image Understanding and Reasoning: A Benchmark and a Training-Free Multi-Agent System

UAV-MAS is proposed, a training-free multi-agent system for MLLM-based UAV aerial image understanding and reasoning, comprising a Domain-Specific Perception Engine that routes queries to task-appropriate visual tools, a Context-Aware Iterative Refinement module (CAIR) that validates intermediate reasoning to curb error accumulation, and a Difficulty-Aware Adaptive Search mechanism (DAAS) that adjusts search depth to question difficulty.

Haoyu Zhang, Shuoxun Zhang, Peng Ye et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.