Jul 2026· ACM Transactions on Software Engineering and Methodology· 0 citations· 53 references
TL;DR
This paper proposes a run-time quantitative operational monitoring methodology based on Discrete-Time Markov Chains (DTMCs), integrating execution traces with probabilistic model checking, and advances agent verification from pre-deployment analysis to ongoing quantitative assurance under uncertainty.
Abstract
The BDI (Belief-Desire-Intention) agent paradigm has been used to model decision-making in autonomous systems that operate with little to no human intervention, reasoning with beliefs, desires, and intentions on-the-fly to achieve goals. Since BDI agents are used to make decisions in safety-critical systems, verifying that they make correct decisions is vital. Existing design-time verification techniques are effective before deployment, but they cannot capture how agent behaviour evolves during execution, especially under imprecise actuation and probabilistic decision-making. This paper bridges the gap between design-time and run-time verification of probabilistic BDI agents. We propose a run-time quantitative operational monitoring methodology based on Discrete-Time Markov Chains (DTMCs), integrating execution traces with probabilistic model checking. Our approach constructs a DTMC representation of possible agent behaviours at design time, then incrementally updates it using observed execution traces to provide situational insight into agent operation. To demonstrate our work, we contribute a probabilistic BDI language extension, explicate our approach through smart manufacturing for probabilistic BDI agents, and use a rover case study to show the generality of our DTMC-based monitoring. We further implement this monitoring mechanism and evaluate its efficiency and scalability experimentally. This advances agent verification from pre-deployment analysis to ongoing quantitative assurance under uncertainty.
We propose a formal model for testing AI agents in stateful environments, using Finite State Machine (FSM) semantics to enable deterministic, trace-level evaluation via an explicit test oracle. Our approach introduces a classified action alphabet and six computable failure-mode detectors for automatic analysis over age...
Ahmed Twabi, Yepeng Ding, Tohru Kondo· International Conference on...· 0 citations
A controlled, physics-grounded benchmark built around planning-induced control trajectories: the ordered planning operations and directives through which an execution architecture acts on other agents and the physical process is introduced.
Results show that a fixed-weight, self-evolving harness can revise, recover, and accumulate verified approaches while producing structured trajectories for future supervised and reinforcement learning.
Boxiu Li, Zi-Mo Wen, Yijia Fan et al.· 2 citations· ⚡1
Requirements engineers for agentic-AI domains face challenges in evaluating, specifying, and operationalizing safe autonomy. Mainstream frameworks, such as Goal-Oriented Requirements Engineering (GORE), lack mechanisms to systematically address these challenges under epistemic uncertainty. We contribute an approach tha...
This work developed AgentInspect, a framework that automatically detects six types of behavioral failures in LangChain-based AI agents by analyzing their execution trajectories across three evaluation settings: a baseline setting using real tool responses, a simulated setting incorporating synthetic tool responses, and...
Ruchira Manke, Mohammad Wardat, Hridesh Rajan et al.· 1 citation
This survey synthesizes 257 papers spanning agent evaluation, software assurance, cyber-physical systems, runtime monitoring, and regulatory guidance in order to characterize the validation problem for agentic systems, and concludes with a lifecycle-oriented research agenda centered on bounded-autonomy specifications,...
Fabio Orazio Mirto, L. D'Agati, Giuseppe Tricomi et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.