Large language model agents are increasingly deployed for long-horizon task execution, raising a central granularity question for trajectory evaluation: whole-trajectory verification is too coarse to capture concrete failures and their associated evidence in long trajectories, while atomic-step scoring is too fine-grai...
Zhichao Shi, Xuhui Jiang, Wen-Jie Zhang et al.· 1 citation
Envs-FORGE is presented, a prompting policy that converts verifier rewards into per-seed environment-synthesis actions and solves a per-seed mixed-integer linear program (MILP) to choose the action that conditions generation.
Xiao-Jun Wu, Ce-Hao Yang, Hong-Hao Liu et al.· 2 citations
The DataFoundry is introduced, a framework for evolving data preparators through recursive self-improvement before large-scale data production, and it is found that recursively evolved preparators produce training data with higher downstream utility than baselines.
Ce-Hao Yang, Xiao-Jun Wu, Xueyuan Lin et al.· 1 citation
LazyTrain is proposed, an optimization layer over a layer-streaming executor that formulates checkpoint selection, activation placement, recomputation, and CPU-GPU-NVMe communication overlap as a mixed-integer scheduling problem, then executes the solved policy during training.
Xiao-Jun Wu, Ce-Hao Yang, Hong-Hao Liu et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.