Jul 2026· 55th International Conference on Environmental Systems· 0 citations
TL;DR
EVA-Bench is presented, a benchmark designed to evaluate foundation model capabilities for exploration EVA (xEVA) support under operationally grounded and safety-relevant conditions and helps identify which models are most suitable for later integration into EVA decision-support systems.
Abstract
Future exploration EVA operations, especially Mars surface EVAs and some higher-tempo Artemis scenarios, will require greater crew autonomy than the International Space Station paradigm because of communication latency, limited bandwidth, and increased operational complexity. Although large language models (LLMs) and agentic AI systems show promise as onboard decision-support tools, their suitability for safety-critical EVA operations remains unclear due to the lack of domain-specific evaluation frameworks. This paper presents EVA-Bench, a benchmark designed to evaluate foundation model capabilities for exploration EVA (xEVA) support under operationally grounded and safety-relevant conditions. EVA-Bench comprises 651 tasks across six EVA scenario families, three difficulty tiers, and two complementary tracks: Single-Query (SQ) tasks for knowledge retrieval, procedural reasoning, and evidence attribution, and End-to-End (E2E) tasks for multi-step planning, tool use, replanning, and protocol compliance in dynamic mission scenarios. Tasks are grounded in a curated corpus of 84 NASA documents spanning Apollo, ISS, Artemis, EVA standards, and mishap investigations. To jointly assess capability and operational safety, the benchmark integrates a safety-sentinel framework informed by Systems-Theoretic Process Analysis, where critical protocol violations zero the final score regardless of task quality. We evaluate nine models from OpenAI, Google, and Anthropic. Results show that models perform strongly on procedural knowledge retrieval, with top SQ scores above 0.93, but degrade substantially on E2E agentic execution, where multi-step planning and contingency handling remain challenging. Results also show that smaller or mid-tier models can outperform larger models on this domain-specific benchmark, suggesting that targeted training and reasoning design may matter more than model scale alone for safety-critical EVA support. Several models with strong objective performance also trigger safety-critical violations, underscoring that raw capability alone is insufficient for operational deployment. By isolating foundation-model capability from higher-level agentic workflow design, EVA-Bench helps identify which models are most suitable for later integration into EVA decision-support systems.
As LLMs increasingly act through tools, they must reconcile user instructions, parametric knowledge, and dynamic environmental observations before taking actions. We introduce KC-Bench, a controlled multi-turn benchmark for measuring this capability across world-knowledge conflicts, input inconsistencies, and multi-source temporal conflicts. Its 238 tasks are manually screened from more than 1,000 generated candidates and combine a user simulator, stateful tools, deterministic environment assertions, an open-source natural-language evaluator, and human trajectory verification. Evaluation of nine models, including DeepSeek-V4-Flash, GLM-5.2, and MiniMax-M3, shows substantial cross-domain variation: no model handles factual correction, identity consistency checking, and temporal conflict resolution reliably across all settings. In the simulated environments, missed conflicts can propagate to tool calls or synthetic protected-data flows. KC-Bench isolates this model-level behavior rather than ranking complete agent frameworks, and provides a reproducible diagnostic for developing conflict-aware reasoning and execution safeguards.
Yaxing Lyu, Sheng-Jie Zhou, B. Toh et al.· 0 citations
EV-WM represents candidate quality with feature and event scores, but these scores do not explicitly record an unmet task predicate, a route label for an available correction mechanism, or a post-correction acceptance result. We present Onto-EV-WM, an ontology-grounded diagnosis and verification-gated correction interface layered above EV-WM rather than a replacement world-model architecture. The implemented task-local TBox defines entity types, predicate signatures, and constraints; source-specific grounding maps predicted or simulator-observed states to task ABoxes; and deterministic rules retain each missing predicate and its arguments when assigning a route label. Learned or heuristic proposers remain separate from this symbolic interface; native task predicates determine acceptance, and the bounded protocol determines whether a failed verification is retried. In the aligned PointMaze evaluation, EV-WM and Onto-EV-WM both report 94% success, with mean final-state distances of 0.90573 and 0.61177, respectively; the separately budgeted search reaches 100% success. On LIBERO-Goal, the ontology represents failed task conditions as typed records, retains their predicate arguments, and associates them with the declared source/joint correction route and predicate-gated acceptance; the complete configuration reports 93.8% corrected-window success on seed 0 and 94.05 +- 0.30% across four evaluation-sampling seeds. On the fixed 10,030-task LIBERO-Plus registry, Onto-EV-WM succeeds on 8,526 tasks (85.00%), with suite-level success rates of 65.98% for LIBERO-10, 91.39% for LIBERO-Goal, and 91.38% for both LIBERO-Object and LIBERO-Spatial. These numbers report the performance of the complete ontology-grounded configurations under the tested simulator protocols; an ontology-only causal share is not measured separately, and real-robot recovery is not evaluated.
Kailin Wang, Haoxiang Jie, Yaoyuan Yan et al.· 0 citations
The results suggest that robust Agentic EDA requires not only stronger models but also structured tool interfaces, persistent design context, controlled execution, and process-level evaluation.
SQBench is introduced, a benchmark for evaluating production-oriented task delivery by language-model agents and shows that functional completion alone does not fully characterize delivery quality and that risk determinations should be reported separately.
DSAgentBench is introduced, the first benchmark to evaluate whether agents can automate full data-science workflows inside real computer environments, and reveals a substantial capability gap between current agentic systems and real data-science workflows.
Mizanur Rahman, Mohammed Saidul Islam, Ridwan Mahbub et al.· 0 citations
This work formalizes minimum-sufficient execution and the Agent Cognitive Redundancy Ratio (ACRR), and proposes E3 (Estimate, Execute, Expand): the agent estimates an initial operating point, executes a minimum viable path, and expands scope only when verification fails.
J. Yin, Xinyu Feng· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.