Skip to content

CurateEvo: Data-Curation Evolving for Agentic Post-Training

Jul 2026 · arXiv.org · Vol abs/2607.06140 · 0 citations · 41 references
Computer Science

TL;DR

CurateEvo, a failure-driven dynamic evolution framework for agentic post-training data curation, which represents the curation strategy as executable code and iteratively rewrites it using failed trajectories from a held-out development set and substantially reduces curation overhead.

Abstract

Large language model (LLM) agents require post-training methods that can improve long-horizon decision making from environment feedback. However, existing agentic post-training pipelines often treat data curation as a fixed preprocessing step, focusing mainly on data augmentation while neglecting filtering, refinement, and adaptation to downstream failures. We propose CurateEvo, a failure-driven dynamic evolution framework for agentic post-training data curation. CurateEvo represents the curation strategy as executable code and iteratively rewrites it using failed trajectories from a held-out development set. At each epoch, the evolved strategy transforms a fixed raw corpus into supervised fine-tuning data, reinforcement learning data, and an inference-time memory bank. The evolution process first improves effectiveness by diagnosing recurring failure modes and augmenting, filtering, or refining data accordingly, and then improves efficiency by pruning redundant or low-utility training turns under a cost-aware objective. Experiments on ACEBench-Agent, BFCL-V4, and {\tau}^2-Bench under both labeled and wild-data settings show that CurateEvo consistently outperforms prior curation methods, improving average scores by 3.2 and 2.7 points, respectively. Further analyses demonstrate that CurateEvo is compatible with different post-training recipes and substantially reduces curation overhead.

View source

Similar papers

Preprint Aug 2026

ADE: Agentic Data Evolution Framework for Human-Centered Objectives

Agentic Data Evolution is proposed, a data-centric framework that organizes synthetic supervision as evolving data snapshots through a closed-loop Observation-Variation-Selection procedure, where a steady-state admission mechanism acts as a quality ratchet that conservatively gates updates for sustained cross-round imp...

Yang Yu, Yi-Lin Jiang, Zexuan Fei et al. · 0 citations
#artificial intelligence Preprint Aug 2026

ReToolSQL: Agentic Reinforcement Learning for Robust Text-to-SQL

ReToolSQL is presented, a two-stage training framework for text-to-SQL that combines a supervised warm-start on rejection-sampled reasoning traces with agentic reinforcement fine-tuning (RFT) over multi-turn tool-use trajectories and shows that a properly designed SFT$\to-RFT pipeline over tool-use trajectories is a pr...

Pratik Kakkar, Chandra Dhir, Ravi Shankar et al. · 0 citations

Rethinking Self-Evolution: A Constrained Exploration-Exploitation Process for Mitigating Skill Overfitting

SkillBoost is proposed, a three-stage framework that mitigates both risks: structured exploitation localizes observed failures to editable skill components, prior-guided exploration draws on prior knowledge in the LLM to generate diverse repair candidates, and verified acceptance commits a candidate only when it improv...

Hong-Qiang Lin, Chao Liu, Xiaofan Bai et al. · 1 citation
Jul 2026

Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering

Recursive self-improvement (RSI) requires AI systems that improve the process of building AI (i.e., AI4AI); machine learning engineering (MLE) offers a concrete, executable testbed for studying this capability. We introduce OpenMLE, an open full-stack system for RSI research in MLE, spanning verifiable task environment...

Junlin Yang, Che Jiang, Yu Fu et al. · 3 citations
#natural language process... Preprint Aug 2026

CAST: Critique-Aware Supervision for Training Reliable Long-Horizon Tool-Calling Agents

CAST, a critique-aware training framework that converts sparse task outcomes into action-level supervision for critique learning and policy optimization, is presented, demonstrating that critique-aware training improves the robustness of LLM agents in realistic dynamic environments.

Amir Saeidi, Zeng Zhang, Rishi Singh et al. · 1 citation
#machine learning Preprint Sep 2026

DE-Venus: A Data-Efficient RLVR Framework for Large Language Models

Reinforcement learning with verifiable rewards (RLVR) improves large language model reasoning, but its practical scaling is constrained by expensive on-policy rollouts and the cost of obtaining reliable targets at scale. Existing methods address sample selection, incomplete supervision, or noisy labels separately, ofte...

Shen-Zhi Yang, Guang-Cheng Zhu, Kai Tang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.