Skip to content

PhysAI-Bench: A Benchmark for LLM-Based Agentic Decision-Making in Autonomous UAV-Centric Physical AI

Sep 2026 · 0 citations · 31 references
Computer Science

TL;DR

The PhysAI-Bench is introduced, a benchmark for evaluating the agentic decision-making required for reliable autonomy in Physical AI, which contains 10,178 standardized decision instances automatically extracted from conversational traces of autonomous UAV missions.

Abstract

Recent advances in Physical AI have accelerated the use of foundation models in autonomous systems such as unmanned aerial vehicles (UAVs), which must perceive, reason, plan, and act in dynamic environments. Existing benchmarks assess physical perception, intuitive physics, embodied navigation, and collaborative reasoning, but rarely evaluate the agentic decision-making required for reliable autonomy. We introduce \textit{PhysAI-Bench}, a benchmark for evaluating this capability. It contains 10,178 standardized decision instances automatically extracted from conversational traces of autonomous UAV missions. Each instance preserves mission context, temporal dependencies, physical constraints, Model Context Protocol (MCP) tool calls, Agent-to-Agent (A2A) interactions, sensor observations, and AI-native 6G network conditions, including latency, packet loss, throughput, edge load, and network slicing. We expose only information preceding each decision, preventing future-event leakage and approximating online decision-making. We evaluate 29 foundation models using a two-stage protocol. We select model-specific configurations from 12 combinations of zero-, three-, and five-shot prompting and four temperatures, tested in three runs on a 35-instance, human-verified development set. We then freeze each selected configuration and evaluate it in three runs on a fixed, episode-disjoint set of 500 instances. GPT-5.3 achieves the highest accuracy (52.00%), followed by GPT-5.2 (49.40%) and Grok~4.5 (49.07%). Few-shot prompting generally improves performance, while temperature has limited influence. The results demonstrate that reliable agentic decision-making in Physical AI remains an open challenge. The dataset is available at https://github.com/maferrag/physai-bench

View source

Similar papers

Open access Nov 2025

UAVBench: An Open Benchmark Dataset for Autonomous and Agentic AI UAV Systems via LLM-Generated Flight Scenarios

Autonomous aerial systems increasingly rely on large language models (LLMs) for mission planning, perception, and decision-making; yet, the lack of standardized, physically grounded benchmarks limits systematic evaluation of their reasoning capabilities. To address this gap, we introduce UAVBench, an open benchmark dat...

M. Ferrag, Abderrahmane Lakas, Mérouane Debbah · 16 citations
#artificial intelligence Preprint Sep 2026

Neuro-Symbolic Agentic AI for Networked Low-Altitude UAVs

Networked low-altitude unmanned aerial vehicles (UAVs) need reliable and adaptive decision-making capabilities to operate under uncertain observations, dynamic environments, and intermittent connectivity, while many existing agentic systems remain limited by hallucination risks, data dependence, and weak generalization...

Yu-Qi Ping, Tian-Hao Liang, Nan-Chi Su et al. · 0 citations
Preprint Aug 2026

LUCID: An Agentic AI Framework on Digital-Twin in the Loop for QoS-Guaranteeing Robotic Control

LUCID is presented, an LLM-agent--orchestrated, uplink-aware cloud-robotics pipeline that moves TP--RRM from solving a fixed formulation to dynamically orchestrating optimization problem schemas within a DITL environment and robustly adapts to changing intents, active-robot counts, and scenes.

Hyeonsu Lyu, Minwoo Kim, Sehyun Ryu et al. · 0 citations
#edge computing Review Open access Sep 2026

Large Language Model-Driven Autonomous UAV Systems: Technical Evolution, Core Architectures, and Critical Challenges

The rapid development of large language models (LLMs) has expanded the capabilities of autonomous unmanned aerial vehicle (UAV) systems in naturallanguage instruction understanding, multimodal perception, and decision-making. This survey reviews the technical evolution, system architectures, and deployment challenges o...

Mei-Jie Zhang, Hao Wang · 0 citations
Open access 2026

Enabling Consensus for Agentic Cooperative Decision-Making in Vehicular Networks

Recent advances in Artificial Intelligence (AI), especially agentic AI, are pushing distributed autonomous systems beyond predefined rule execution toward joint decision-making among autonomous participants with potentially different models, preferences, constraints, and value assessments. In such environments, coordin...

Huan-Yu Wu, Chen-Tao Yue, Yan Gou et al. · 0 citations
#artificial intelligence Preprint Sep 2026

A Simulation Platform for AUV Fault Recovery: Exploring LLM-Based Diagnostic Strategies

Autonomous underwater vehicles (AUVs) operating beyond reliable communications must recover from failures without human intervention. We investigate an architecture in which conventional deterministic layered control autonomy manages normal operations, while an invokable large language model (LLM) serves as a diagnosti...

Khalid Halba, Kylie Cooper, James G. Bellingham · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.