Skip to content

Beyond Endpoint Performance: Process-Level Evaluation of Self-Evolving Agents

Sep 2026 · 1 citation
Computer Science

TL;DR

EvoPathBench establishes capability-level process evaluation as a foundation for analyzing self-evolution, identifying candidate evaluation and selection as key targets for improvement.

Abstract

Self-evolving agents convert interaction feedback into persistent artifacts, such as memories or skills, which in turn guide subsequent decisions. As these artifacts are iteratively updated throughout an experience stream, the capabilities they support may evolve. Consequently, endpoint performance alone offers an incomplete view of self-evolution. Process-level evaluation is therefore essential to identify when a target capability emerges and whether later updates strengthen, preserve, or weaken it. Motivated by this, we propose \textsc{EvoPathBench}, a benchmark that tracks individual capabilities during artifact-level self-evolution. EvoPathBench fixes the base model, tools, freezes evolving artifacts at successive checkpoints, and evaluates the target capability on held-out episodes. This benchmark evaluates agent self-evolution using public trading data and calibrated trajectories. It tests three capabilities: generalization to unseen tasks, retention after unrelated learning, and rule adaptation to new evidence. Experimental results show that gains on similar unseen tasks often weaken under distribution shift, retention losses are concentrated in a minority of evolution paths, and no method achieves reliable rule adaptation. Moreover, while self-evolution enables agents to generate candidate artifacts with substantial held-out gains, the selected updates consistently fall short of realizing this potential. Together, these findings establish capability-level process evaluation as a foundation for analyzing self-evolution, identifying candidate evaluation and selection as key targets for improvement.

View source

Similar papers

#artificial intelligence Preprint Oct 2026

Capability-Driven Self-Evolution of Agent Memory

Memory self-evolution uses task feedback to iteratively improve executable memory programs that store and retrieve information from past interactions. Existing approaches typically adopt holistic evolution, deriving revision directions from mixed feedback and judging progress by overall performance. This can obscure op...

Yao-Qi Chen, Yu-Ru Feng, Qianxi Zhang et al. · 0 citations
#natural language process... Preprint Sep 2026

Experience Funnel: A State-Policy Alternating Loop for Self-Evolving Agents

Autonomous agents powered by large language models (LLMs) continuously accumulate experience through interaction, creating an opportunity to improve future behavior through self-evolution. A fundamental challenge is how to transform abundant, task-specific interaction experience into reusable model competence without s...

Wen-Bo Gao, Zhao-Mou Song, Zhi-Yuan Ji et al. · 0 citations
Preprint Aug 2026

Tree-of-Experience: Hierarchical Experience Management for Self-Evolving Agents

Experimental results show that ToE substantially improves both problem-solving performance and efficiency and organizes the experience into a shared tree of analytical perspectives and reasoning paths, whose reliability is calibrated through environmental outcomes to support systematic updating, transfer, and efficient...

Zi-Hao Deng, Yi-Ning Zhu, Lei-Ming Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

HarnessEvolve: Learning from Reference Trajectories for Reliable Agent Self-Evolution

This paper introduces HarnessEvolve, a self-evolving framework that learns from reference trajectories to achieve reliable agent self-evolution, and overcomes credit assignment failure by generating reference trajectories and aligning failed executions against them to extract error signals.

Wen Jiang, Ming-Min Chu, Yiding Tian et al. · 7 citations
#artificial intelligence Preprint Sep 2026

When Successful Strategies Fail: Adaptation to Environmental Novelty in Terminal Agents

This work introduces AGNI, an automated pipeline that extracts trajectory-relevant assumptions, injects targeted environmental changes, and validates that the resulting novel tasks remain solvable and highlights a gap between task competence and adaptive capability and motivate environmental variation as a core dimensi...

Janvijay Singh, Vaishnavi Shrivastava, Dilek Hakkani-Tur et al. · 0 citations
#natural language process... Preprint Sep 2026

The Right Lesson at the Right Step: Deriving Control Updates for Self-Evolving Agents

Self-evolving agents improve future behavior by reusing past experience, typically as global prompts, memories, or reflections. Yet these mechanisms rarely control where experience takes effect. In long tool-use workflows, the same lesson may correct one decision but distract another, making experience reuse a problem...

Yun-He Su, Zi-Yi Dong, Tong Yu et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.