Formal Trajectory Analysis for Testing Agentic AI in Stateful Environments
We propose a formal model for testing AI agents in stateful environments, using Finite State Machine (FSM) semantics to enable deterministic, trace-level evaluation via an explicit test oracle. Our approach introduces a classified action alphabet and six computable failure-mode detectors for automatic analysis over agent traces. Applying this methodology to 350 multi-turn network configuration runs, we find that aggregate scores often miss structural failures revealed by trajectory analysis, and that meltdown rates alone fail to distinguish distinct failure mechanisms. Agents satisfy locally verifiable intent properties more than twice as often as remotely verifiable ones, exposing a critical awareness–action gap (the difference in satisfaction rate between local and remote evidence). Trajectory-level analysis is therefore essential for robust evaluation of autonomous agents in complex systems.