Skip to content

Policy Drift in Learning AI Agents: A Dynamical Systems Perspective on Security Degradation

2025 · International Journal of Scientific Research and Management · Vol 13, pp. 2457-2463 · 1 citation

TL;DR

Policy drift takes shape through a nonlinear differential equation - framed within the policy state space - with support from Lyapunov stability concepts alongside bifurcation methods alongside bifurcation methods, and the Intent Drift Rate appears: a concrete number per dialogue turn built as the time-based change in the Behavioral Drift Index.

Abstract

Paper 1 of this series introduced clean attacks adversarial inputs that satisfy all observable policy constraints yet redirect autonomous agent behavior away from operator intent. Still, uncertainty remains about how such shifts unfold across extended usage. Even though immediate effects suggest weakened breaches cause less impact (58.1%) compared to full goal swaps (64.8%), real-world cases often describe worsening divergence after multiple cycles, hinting at gradual breakdowns needing deeper modeling. Here, deviation over time emerges as slow but compounding movement of an agent’s actual choices beyond intended boundaries Π₀ due to recurring subtle assaults. This shift follows complex patterns - not linear decay, rather like systems building hidden strain until sudden collapse occurs past threshold t, then veering sharply into attacker-influenced regimes Π without return. Beyond tipping point λ in frequency of these disguised interventions, recovery fades; alignment dissolves not by force, but through accumulation. This study offers four core advances. Right away, policy drift takes shape through a nonlinear differential equation - framed within the policy state space - with support from Lyapunov stability concepts alongside bifurcation methods. Instead of vague metrics, the Intent Drift Rate appears: a concrete number per dialogue turn, built as the time-based change in the Behavioral Drift Index first laid out in Paper 1. Following that, Π₀’s resilience gets examined via a two-valley energy model, uncovering how the tipping point labeled λ controls whether systems rebound or shift into Π. Later still, validation arrives through tests on AegisBench-MT - a version expanded to 480 conversations - spanning three types of agents plus four leading large language models, spotting clear shifts near t ≈ 9 steps along with a split triggered around λ ≈ 0.42. Foundational input comes from IDR, shaping core logic across Papers 3 through 5. Trust decisions driven by user purpose rely on it, just as node-level shift tracking within agent networks does. Response mechanisms that activate when deviations occur also build upon this base. Each test applies common reference terms defined earlier in Paper 1 - elements like φ, Π, SVS, and BDI remain consistent throughout.

Read PDF

Similar papers

#artificial intelligence Preprint Sep 2026

Representation Transitions Reveal Emerging Safety Risks in Multi-Turn LLM Agents

Multi-turn attacks on agentic systems can compose individually permissible actions into harmful outcomes, challenging defenses that assess actions or states in isolation. We show that such attacks leave a detectable signature in the agent's internal representations: harmful behavior emerges as an accumulated representa...

Hao-Yu Wang, Wei Zhao, Ye-Di Zhang et al. · 0 citations
#machine learning Preprint Oct 2026

FoSeRL: Formal Sequential Robustness Certification for Reinforcement Learning Policies

Even a few action perturbations can substantially degrade the performance of a deployed decision policy. Certifying the resulting return loss is challenging in stochastic environments, where returns vary even without an attack. We introduce FoSeRL, a framework for certifying deployed RL policies against precommitted, t...

Sara Taheri, D. Ganguly, Jan Křetinský et al. · 0 citations
Conference Sep 2026

Preventing Control Loss in Stochastic Environments: Recurrent Proximal Policy Optimization with Conditional Value at Risk

Preventing transient control loss in risk-critical human-machine systems requires dynamic intervention strategies that account for unobservable user states. Traditional control frameworks relying on fully observable Markov decision processes are inadequate for this task, as they inherently suffer from perceptual aliasi...

Oleksandr Chaban, V. Hladun · 0 citations
#artificial intelligence Preprint Sep 2026

SoK: Trading Agents or Market Crashers? Dissecting Robustness and Security Failures in Academic Financial LLM Trading Schemes

FARSIGHT (Financial Agent Robustness and Security Investigation and Global Holistic Testing), a framework that performs scheme-level evaluation of financial LLM agents on two axes: robustness under market turbulence and security against three attack types: attacks on information sources, attacks on agents, and agent-as...

Meng-Xiao Wang, Nitesh Saxena · 0 citations
#generative ai Review Open access Sep 2026

From Herding Machines to Autonomous Agents: A Taxonomy of AI-Driven Flash Crash Mechanisms and the Regulatory Gap

A prospective review proposes a three-category taxonomy of the mechanisms through which AI triggers catastrophic, self-reinforcing market dislocations, or “flash crashes,” and concludes with policy recommendations on model-diversity mandates, real-time AI trading surveillance, adaptive circuit-breaker design, and cross...

Kan-Ching Ng · 0 citations

Related blog posts

GPT-Lab Aug 28, 2026

We built an AI factory for HVAC control

What does it take to trust AI-driven HVAC optimization? Our AI Model Factory combines agents, machine learning, reinforcement learning and deterministic checks in a governed workflow designed for messy, real-world building data. The post We built an AI factory for HVAC control appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.