CurateEvo, a failure-driven dynamic evolution framework for agentic post-training data curation, which represents the curation strategy as executable code and iteratively rewrites it using failed trajectories from a held-out development set and substantially reduces curation overhead.
Abstract
Large language model (LLM) agents require post-training methods that can improve long-horizon decision making from environment feedback. However, existing agentic post-training pipelines often treat data curation as a fixed preprocessing step, focusing mainly on data augmentation while neglecting filtering, refinement, and adaptation to downstream failures. We propose CurateEvo, a failure-driven dynamic evolution framework for agentic post-training data curation. CurateEvo represents the curation strategy as executable code and iteratively rewrites it using failed trajectories from a held-out development set. At each epoch, the evolved strategy transforms a fixed raw corpus into supervised fine-tuning data, reinforcement learning data, and an inference-time memory bank. The evolution process first improves effectiveness by diagnosing recurring failure modes and augmenting, filtering, or refining data accordingly, and then improves efficiency by pruning redundant or low-utility training turns under a cost-aware objective. Experiments on ACEBench-Agent, BFCL-V4, and {\tau}^2-Bench under both labeled and wild-data settings show that CurateEvo consistently outperforms prior curation methods, improving average scores by 3.2 and 2.7 points, respectively. Further analyses demonstrate that CurateEvo is compatible with different post-training recipes and substantially reduces curation overhead.
Agentic Data Evolution is proposed, a data-centric framework that organizes synthetic supervision as evolving data snapshots through a closed-loop Observation-Variation-Selection procedure, where a steady-state admission mechanism acts as a quality ratchet that conservatively gates updates for sustained cross-round imp...
Yang Yu, Yi-Lin Jiang, Zexuan Fei et al.· 0 citations
ReToolSQL is presented, a two-stage training framework for text-to-SQL that combines a supervised warm-start on rejection-sampled reasoning traces with agentic reinforcement fine-tuning (RFT) over multi-turn tool-use trajectories and shows that a properly designed SFT$\to-RFT pipeline over tool-use trajectories is a pr...
Pratik Kakkar, Chandra Dhir, Ravi Shankar et al.· 0 citations
SkillBoost is proposed, a three-stage framework that mitigates both risks: structured exploitation localizes observed failures to editable skill components, prior-guided exploration draws on prior knowledge in the LLM to generate diverse repair candidates, and verified acceptance commits a candidate only when it improv...
Hong-Qiang Lin, Chao Liu, Xiaofan Bai et al.· arXiv.org· 1 citation
Recursive self-improvement (RSI) requires AI systems that improve the process of building AI (i.e., AI4AI); machine learning engineering (MLE) offers a concrete, executable testbed for studying this capability. We introduce OpenMLE, an open full-stack system for RSI research in MLE, spanning verifiable task environment...
Junlin Yang, Che Jiang, Yu Fu et al.· arXiv.org· 3 citations
CAST, a critique-aware training framework that converts sparse task outcomes into action-level supervision for critique learning and policy optimization, is presented, demonstrating that critique-aware training improves the robustness of LLM agents in realistic dynamic environments.
Amir Saeidi, Zeng Zhang, Rishi Singh et al.· 1 citation
Reinforcement learning with verifiable rewards (RLVR) improves large language model reasoning, but its practical scaling is constrained by expensive on-policy rollouts and the cost of obtaining reliable targets at scale. Existing methods address sample selection, incomplete supervision, or noisy labels separately, ofte...
Shen-Zhi Yang, Guang-Cheng Zhu, Kai Tang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.