Skip to content

Self-Evolving Agents via Likelihood-Guided Tool-Space Optimization

Sep 2026 · 0 citations · 59 references
Computer Science

TL;DR

This work introduces LOTS (Likelihood-Only Tool Scoring), which evolves an agent's tool space from accumulated output experience while keeping model parameters fixed and demonstrates that task-specific spaces persist and continue to improve over time, while cross-model experiments show that learned configurations transfer across different models.

Abstract

Self-evolving agents can continually improve their behavior, while tools define the executable action space through which they interact with the environment. However, exposing the full tool library to model introduces substantial irrelevant context and can impair tool-use decisions. We study tool-space self-evolution, where each recurring task type maintains a persistent tool space which is constructed from accumulated output experience. We identify three limitations of existing methods: (1) output-unaware selection: they rely primarily on tool descriptions or model priors rather than observed tool outputs; (2) statelessness across request: they select tools independently for each request without consolidating prior output experience into persistent task-specific state; (3) inference cost: they repeatedly search, rank, or reason over candidate tools for subsequent requests of the same task. We address these limitations through output-aware tool scoring, persistent task-specific tool spaces, amortized tool selection, and reusable configurations across models. We introduce LOTS (Likelihood-Only Tool Scoring), which evolves an agent's tool space from accumulated output experience while keeping model parameters fixed. After each request, LOTS holds the model's generated answer and estimates each tool's contribution by measuring how much the answer likelihood changes when its observed output is removed. These contributions are aggregated within each recurring task to rank tools and update its persistent space. Across three benchmarks, LOTS improves task performance while substantially reducing tool context. More importantly, sequential experiments demonstrate that task-specific spaces persist and continue to improve over time, while cross-model experiments show that learned configurations transfer across different models.

View source

Similar papers

#natural language process... Preprint Sep 2026

Experience Funnel: A State-Policy Alternating Loop for Self-Evolving Agents

Autonomous agents powered by large language models (LLMs) continuously accumulate experience through interaction, creating an opportunity to improve future behavior through self-evolution. A fundamental challenge is how to transform abundant, task-specific interaction experience into reusable model competence without s...

Wen-Bo Gao, Zhao-Mou Song, Zhi-Yuan Ji et al. · 0 citations
Preprint Aug 2026

Rethinking Self-Evolving Agents: Do We Still Need Prescribed Optimization Pipelines?

This work introduces Open-Ended Optimization (OEO), which keeps the objective, permitted interactions, resource budget, data boundary, and evaluation fixed while allowing the optimizer to compose the improvement process online.

Xue Hui, Fan Yang · 3 citations
#artificial intelligence Preprint Sep 2026

When Successful Strategies Fail: Adaptation to Environmental Novelty in Terminal Agents

LLM agents increasingly solve long-horizon tasks by autonomously interacting with their environment. In doing so, their strategies rely on assumptions about that environment: which resources and tools exist, where they are located, and how they behave. When these assumptions no longer hold, reliable agents must detect...

Janvijay Singh, Vaishnavi Shrivastava, Dilek Hakkani-Tur et al. · 0 citations
#natural language process... Preprint Sep 2026

The Right Lesson at the Right Step: Deriving Control Updates for Self-Evolving Agents

Self-evolving agents improve future behavior by reusing past experience, typically as global prompts, memories, or reflections. Yet these mechanisms rarely control where experience takes effect. In long tool-use workflows, the same lesson may correct one decision but distract another, making experience reuse a problem...

Yun-He Su, Zi-Yi Dong, Tong Yu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

CompoWorld: Compositional Environment Scaling for General Agents

Automatically generated environments provide a scalable source of interaction data for training general agents. However, existing approaches mainly generate tasks within a single environment, while real-world workflows require agents to connect information and actions across multiple services. We introduce Compositiona...

Xiao-Wen Yang, Wei-Yi Xu, Wen Da et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SelfSearch: Reward-Free Search for Self-Improving Agents

Advances in the coding capabilities of LLM agents allow them to inspect and modify their own instructions, tools, and execution procedures. Existing approaches use this ability to search for improved agents through repeated downstream evaluation, which incurs substantial costs and ties the search to the evaluated tasks...

Jungwoo Yang, InJin Kong, Yohan Jo · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.