2025· International Journal of Scientific Research and Management· Vol 13, pp. 2457-2463· 1 citation
TL;DR
Policy drift takes shape through a nonlinear differential equation - framed within the policy state space - with support from Lyapunov stability concepts alongside bifurcation methods alongside bifurcation methods, and the Intent Drift Rate appears: a concrete number per dialogue turn built as the time-based change in the Behavioral Drift Index.
Abstract
Paper 1 of this series introduced clean attacks adversarial inputs that satisfy all observable policy constraints yet redirect autonomous agent behavior away from operator intent. Still, uncertainty remains about how such shifts unfold across extended usage. Even though immediate effects suggest weakened breaches cause less impact (58.1%) compared to full goal swaps (64.8%), real-world cases often describe worsening divergence after multiple cycles, hinting at gradual breakdowns needing deeper modeling. Here, deviation over time emerges as slow but compounding movement of an agent’s actual choices beyond intended boundaries Π₀ due to recurring subtle assaults. This shift follows complex patterns - not linear decay, rather like systems building hidden strain until sudden collapse occurs past threshold t, then veering sharply into attacker-influenced regimes Π without return. Beyond tipping point λ in frequency of these disguised interventions, recovery fades; alignment dissolves not by force, but through accumulation. This study offers four core advances. Right away, policy drift takes shape through a nonlinear differential equation - framed within the policy state space - with support from Lyapunov stability concepts alongside bifurcation methods. Instead of vague metrics, the Intent Drift Rate appears: a concrete number per dialogue turn, built as the time-based change in the Behavioral Drift Index first laid out in Paper 1. Following that, Π₀’s resilience gets examined via a two-valley energy model, uncovering how the tipping point labeled λ controls whether systems rebound or shift into Π. Later still, validation arrives through tests on AegisBench-MT - a version expanded to 480 conversations - spanning three types of agents plus four leading large language models, spotting clear shifts near t ≈ 9 steps along with a split triggered around λ ≈ 0.42. Foundational input comes from IDR, shaping core logic across Papers 3 through 5. Trust decisions driven by user purpose rely on it, just as node-level shift tracking within agent networks does. Response mechanisms that activate when deviations occur also build upon this base. Each test applies common reference terms defined earlier in Paper 1 - elements like φ, Π, SVS, and BDI remain consistent throughout.
Multi-turn attacks on agentic systems can compose individually permissible actions into harmful outcomes, challenging defenses that assess actions or states in isolation. We show that such attacks leave a detectable signature in the agent's internal representations: harmful behavior emerges as an accumulated representa...
Hao-Yu Wang, Wei Zhao, Ye-Di Zhang et al.· 0 citations
Even a few action perturbations can substantially degrade the performance of a deployed decision policy. Certifying the resulting return loss is challenging in stochastic environments, where returns vary even without an attack. We introduce FoSeRL, a framework for certifying deployed RL policies against precommitted, t...
Sara Taheri, D. Ganguly, Jan Křetinský et al.· 0 citations
Preventing transient control loss in risk-critical human-machine systems requires dynamic intervention strategies that account for unobservable user states. Traditional control frameworks relying on fully observable Markov decision processes are inadequate for this task, as they inherently suffer from perceptual aliasi...
Oleksandr Chaban, V. Hladun· Automation, Control, and Inf...· 0 citations
FARSIGHT (Financial Agent Robustness and Security Investigation and Global Holistic Testing), a framework that performs scheme-level evaluation of financial LLM agents on two axes: robustness under market turbulence and security against three attack types: attacks on information sources, attacks on agents, and agent-as...
A prospective review proposes a three-category taxonomy of the mechanisms through which AI triggers catastrophic, self-reinforcing market dislocations, or “flash crashes,” and concludes with policy recommendations on model-diversity mandates, real-time AI trading surveillance, adaptive circuit-breaker design, and cross...
Kan-Ching Ng· Journal of Risk and Financia...· 0 citations
What does it take to trust AI-driven HVAC optimization? Our AI Model Factory combines agents, machine learning, reinforcement learning and deterministic checks in a governed workflow designed for messy, real-world building data. The post We built an AI factory for HVAC control appeared first on GPT-Lab.
MIT News · Artificial Intelligence· news.mit.eduAug 18, 2026
A new method for surgically removing training examples from a model reveals that as datasets grow, the link between what a model learns and what it produces dissolves.
MIT News · Artificial Intelligence· news.mit.eduMay 20, 2026