Fault Propagation and Risk Containment in Agentic AI Systems
Abstract
Agentic artificial intelligence systems differ from conventional language model applications because they can reasonover a task, select tools, invoke external systems, observe resulting state, and continue acting toward a goal. Thisaction loop creates a security and reliability problem that cannot be adequately described by model accuracy orprompt-level safety alone. A local reasoning error, malicious instruction, misleading tool output, or authorizationmistake can propagate across subsequent actions and produce a system-level failure with consequences larger thanthe initial fault. This paper proposes a systems framework for analyzing that problem. It introduces Agentic FaultPropagation as the transmission or amplification of a local fault across an agent's perception, reasoning, planning, toolselection, authorization, execution, observation, and replanning cycle. It also introduces Agentic Blast Radius as ameasure of the maximum consequential impact that can result from an undetected fault before containment or humanintervention. Building on established work on tool-using language models, agent evaluation, prompt injection, leastprivilege, and AI risk management, the paper develops a fault taxonomy and a containment architecture. Theframework has four components: a fault lifecycle model, a propagation graph, a blast-radius model, and a layeredcontainment strategy. Proposed controls include least-privilege tool authorization, deterministic policy enforcement,stage-level validation, trust boundaries around external content, independent approval for high-impact actions, runtimemonitoring, provenance-aware logging, and recovery mechanisms. The paper does not claim experimental results.Instead, it specifies an empirical evaluation protocol that can be implemented using agent benchmarks and controlledtool environments. The central research claim is that secure agentic AI should be evaluated not only by whether anagent can complete a task, but also by how far an error can travel before the system stops it.