Skip to content

SAGE: A Statistical Acceptance Gate for Self-Evolving Agents

Sep 2026 · 0 citations · 30 references
Computer Science

TL;DR

SAGE is a conservative refinement of the standard gate that recovers the baseline exactly at a boundary setting, and employs a one-sided paired test that commits an edit only when its wins are statistically reliable against its losses, and it abstains otherwise.

Abstract

Large Language Model (LLM)-based agents increasingly self-evolve by editing a persistent skill document that encodes their workflow, tool-use rules, and decision logic. This loop has two steps, an optimizer that proposes a candidate edit and a gate that accepts or rejects it. Prior work has concentrated on the optimizer, while the gate still follows a naive rule that keeps any edit which improves an aggregate validation score. We show that this rule fails in two ways. First, it admits permanent regressions, since an edit can raise the average while breaking items the skill already solves. Second, it is vulnerable to the Optimizer's Curse, since the best observed score on a finite and noisy validation set is upward biased. To solve the above two limitations, we propose a statistical acceptance gate for self-evolving agents (SAGE). Compared with previous work, SAGE has two contributions. First, SAGE proposes a per-item paired comparison that evaluates the current skill and the edited skill on identical validation items, which exposes regressions that an aggregate score hides and penalizes them asymmetrically. Second, SAGE also employs a one-sided paired test that commits an edit only when its wins are statistically reliable against its losses, and it abstains otherwise. SAGE is a conservative refinement of the standard gate that recovers the baseline exactly at a boundary setting. It commits only a subset of the baseline's edits, filtering out those whose gains are unreliable or purchased by breaking already-solved items. Across five benchmarks and four backbone LLMs under an equal-budget protocol, SAGE lowers the regression rate in 19 of 20 settings and matches the baseline in the remaining one, for example from 36.5% to 0% on LiveMath and from 42.8% to 0% on OfficeQA with DeepSeek-V4. SAGE also attains the highest final score in all 20 settings, raising LiveMath from 34.15 to 48.78.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

TTSE: A Two-Track Online Self-Evolution Framework

This paper proposes TTSE (Two-Track Self-Evolution), a dual-track online self-evolution framework that separates evolving knowledge into FACT (environmental facts, whose reliability is continuously verified through interaction evidence) and TIP (task-conditioned implementation procedures), and broadly compatible with e...

Rui-Min Pei, Yong-Kang Wu, Shang-Yi Zheng et al. · 0 citations
#artificial intelligence Review Sep 2026

RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models

Auto-research agents, LLM systems that propose, implement, train, and evaluate model changes across iterations, promise to automate applied ML's experimental loop. Over long horizons, execution accuracy is a binding constraint: a change can silently leak held-out data, omit normalization, disconnect a gradient, or leav...

Zheng-Yu Chen, Lin-Feng Liu, Hong Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

StateTape: Action-Conditioned Evidence Lifecycle Modeling for Long-Horizon Coding Agents

Despite the recent success of coding agents built on large language models, it remains challenging to run them over long horizons, since every observation is appended to the context and the context grows with each one. History-based maintenance is a common remedy, which masks or summarizes old observations, or prunes w...

Zi-Yang Yu, Liang Zhao, Bo-Wen Zhu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Disclosure-Gated User Simulation for Companion-Agent Evaluation

A disclosure gate conditioning information release on the companion agent's behaviour is answered, with a disclosure gate conditioning information release on the companion agent's behaviour: its state is a ladder of five ordered gates, merged onto three observable depth layers.

Yaozhong Liu, Yu He · 0 citations
Preprint Aug 2026

Rethinking Self-Evolving Agents: Do We Still Need Prescribed Optimization Pipelines?

This work introduces Open-Ended Optimization (OEO), which keeps the objective, permitted interactions, resource budget, data boundary, and evaluation fixed while allowing the optimizer to compose the improvement process online.

Xue Hui, Fan Yang · 3 citations
#artificial intelligence Preprint Sep 2026

Propose, Don't Judge: An Anytime-Valid Referee for LLM Agents That Mine Investment Factors

Language-model agents now run the whole of quantitative factor research: they propose investment factors, backtest them, select the survivors and retire them, and the answer is governed self-evolution: the agent may propose, and a frozen statistical referee that the agent cannot touch must judge.

Bo Qu, Ming-Guang Chen, Li-Cheng Wang · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.