PhysAI-Bench: A Benchmark for LLM-Based Agentic Decision-Making in Autonomous UAV-Centric Physical AI
The PhysAI-Bench is introduced, a benchmark for evaluating the agentic decision-making required for reliable autonomy in Physical AI, which contains 10,178 standardized decision instances automatically extracted from conversational traces of autonomous UAV missions.
M. Ferrag, Mérouane Debbah, Abderrahmane Lakas et al.
· 0 citations