The PhysAI-Bench is introduced, a benchmark for evaluating the agentic decision-making required for reliable autonomy in Physical AI, which contains 10,178 standardized decision instances automatically extracted from conversational traces of autonomous UAV missions.
Abstract
Recent advances in Physical AI have accelerated the use of foundation models in autonomous systems such as unmanned aerial vehicles (UAVs), which must perceive, reason, plan, and act in dynamic environments. Existing benchmarks assess physical perception, intuitive physics, embodied navigation, and collaborative reasoning, but rarely evaluate the agentic decision-making required for reliable autonomy. We introduce \textit{PhysAI-Bench}, a benchmark for evaluating this capability. It contains 10,178 standardized decision instances automatically extracted from conversational traces of autonomous UAV missions. Each instance preserves mission context, temporal dependencies, physical constraints, Model Context Protocol (MCP) tool calls, Agent-to-Agent (A2A) interactions, sensor observations, and AI-native 6G network conditions, including latency, packet loss, throughput, edge load, and network slicing. We expose only information preceding each decision, preventing future-event leakage and approximating online decision-making. We evaluate 29 foundation models using a two-stage protocol. We select model-specific configurations from 12 combinations of zero-, three-, and five-shot prompting and four temperatures, tested in three runs on a 35-instance, human-verified development set. We then freeze each selected configuration and evaluate it in three runs on a fixed, episode-disjoint set of 500 instances. GPT-5.3 achieves the highest accuracy (52.00%), followed by GPT-5.2 (49.40%) and Grok~4.5 (49.07%). Few-shot prompting generally improves performance, while temperature has limited influence. The results demonstrate that reliable agentic decision-making in Physical AI remains an open challenge. The dataset is available at https://github.com/maferrag/physai-bench
Autonomous aerial systems increasingly rely on large language models (LLMs) for mission planning, perception, and decision-making; yet, the lack of standardized, physically grounded benchmarks limits systematic evaluation of their reasoning capabilities. To address this gap, we introduce UAVBench, an open benchmark dat...
M. Ferrag, Abderrahmane Lakas, Mérouane Debbah· IEEE Open Journal of Vehicul...· 16 citations
Networked low-altitude unmanned aerial vehicles (UAVs) need reliable and adaptive decision-making capabilities to operate under uncertain observations, dynamic environments, and intermittent connectivity, while many existing agentic systems remain limited by hallucination risks, data dependence, and weak generalization...
Yu-Qi Ping, Tian-Hao Liang, Nan-Chi Su et al.· 0 citations
LUCID is presented, an LLM-agent--orchestrated, uplink-aware cloud-robotics pipeline that moves TP--RRM from solving a fixed formulation to dynamically orchestrating optimization problem schemas within a DITL environment and robustly adapts to changing intents, active-robot counts, and scenes.
Hyeonsu Lyu, Minwoo Kim, Sehyun Ryu et al.· 0 citations
The rapid development of large language models (LLMs) has expanded the capabilities of autonomous unmanned aerial vehicle (UAV) systems in naturallanguage instruction understanding, multimodal perception, and decision-making. This survey reviews the technical evolution, system architectures, and deployment challenges o...
Mei-Jie Zhang, Hao Wang· Intelligence & Control· 0 citations
Recent advances in Artificial Intelligence (AI), especially agentic AI, are pushing distributed autonomous systems beyond predefined rule execution toward joint decision-making among autonomous participants with potentially different models, preferences, constraints, and value assessments. In such environments, coordin...
Huan-Yu Wu, Chen-Tao Yue, Yan Gou et al.· IEEE Transactions on Cogniti...· 0 citations
Autonomous underwater vehicles (AUVs) operating beyond reliable communications must recover from failures without human intervention. We investigate an architecture in which conventional deterministic layered control autonomy manages normal operations, while an invokable large language model (LLM) serves as a diagnosti...
Khalid Halba, Kylie Cooper, James G. Bellingham· 0 citations
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.