Skip to content

SEABench: Benchmarking Endogenous Misalignment In Self-Evolving Agents

Sep 2026 · 0 citations · 48 references
Computer Science

TL;DR

This work introduces SEABench, a benchmark for studying endogenous misalignment arising from agent self-evolution, with 48 longitudinal task sequences that span multiple evolution surfaces, task domains, and harm types in a rich personal-assistant environment and shows that qualitatively different safety behaviors emerge across evolution surfaces and harm types.

Abstract

Self-evolving LLM agents have gained prominence for their ability to improve after deployment by modifying their harness, including their controller instructions, memory management protocols, and reusable tools and skills, in response to user and environment feedback. However, locally useful updates may persist into later tasks where they produce unsafe behavior, even without direct adversarial influence. To study this risk, we introduce SEABench, a benchmark for studying endogenous misalignment arising from agent self-evolution, with 48 longitudinal task sequences that span multiple evolution surfaces, task domains, and harm types in a rich personal-assistant environment. To account for the stochasticity inherent in agentic operations, we provide an adaptive trajectory discovery pipeline that probes for failures while preserving original task intent and supports causal attribution through paired non-evolving agents and attribution scores. Our evaluation across multiple recent LLMs, evolution surfaces, and harm types reveals that self-evolution indeed increases task completion rates but often at the cost of safety failures that are absent for paired non-evolving baseline agents. We also show that qualitatively different safety behaviors emerge across evolution surfaces and harm types. Further, we show that this divergence in safety behavior is reflected in agents'chain-of-thought reasoning, which yields an effective monitoring strategy that can mitigate unsafe behavior with a low false positive rate.

View source

Similar papers

Preprint Aug 2026

Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents

Self-improving LLM agents convert successful trajectories into persistent cross-task state. An unsafe success can thereby become reusable policy after its triggering input disappears. Skill evolution makes this failure measurable by distilling operational trajectories into executable, transferable, and inspectable proc...

Xutao Mao, Liang Zhao, Xiang Zheng et al. · 2 citations · ⚡1
#artificial intelligence Preprint Sep 2026

HarnessEvolve: Learning from Reference Trajectories for Reliable Agent Self-Evolution

This paper introduces HarnessEvolve, a self-evolving framework that learns from reference trajectories to achieve reliable agent self-evolution, and overcomes credit assignment failure by generating reference trajectories and aligning failed executions against them to extract error signals.

Wen Jiang, Ming-Min Chu, Yiding Tian et al. · 7 citations
#artificial intelligence Preprint Sep 2026

SelfOp: An Optimization Algorithm for Self-Improving Security Agents

LLM agents are increasingly used for security tasks: vulnerability discovery, exploit reproduction, and patch generation. Improving them at the model level demands expert demonstrations or computable rewards, which security tasks rarely offer: traces are costly, failures hard to diagnose, rewards sparse, and non-comput...

Saad Ullah, Yiğitcan Kaya, Christopher Kruegel et al. · 0 citations
#natural language process... Preprint Sep 2026

EvoHarnessBench: Can Your Agents Keep Pace with an Evolving Harness?

Modern LLM-based agents operate through a harness of tools, reusable skills, and specialist agents that shapes what they observe and what they can do. In practice, this harness continually evolves as new capabilities are added. We introduce EvoHarnessBench, a benchmark for evaluating agents under controlled harness evo...

Zi-Xuan Ke, Vaidehi Patil, Hai-Zhou Shi et al. · 1 citation
Preprint Aug 2026

StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments

StarHarness offers a practical way to reduce persistent model-environment mismatch in tool-rich enterprise tasks by stratifying tasks according to baseline failure behavior, separating proposer-visible search tasks from proposer-hidden selection tasks, and reserves held-out tasks for evaluating generalization.

Esakkivel Esakkiraja, D. Akhiyarov, Vikas Yadav et al. · 1 citation
Preprint Aug 2026

EnvHarness: Awakening Static Worlds for Agent Learning

Environment Harness is proposed, a programmable layer of plug-in components that wraps a static environment to reshape its behavior without modifying the underlying logic, enabling continuous, targeted co-evolution of the policy and its environment.

Chengsong Huang, Zifeng Wang, Ru-Jun Han et al. · 11 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.