EvoPathBench establishes capability-level process evaluation as a foundation for analyzing self-evolution, identifying candidate evaluation and selection as key targets for improvement.
Hong-Qiang Lin, Chao Liu, Xiao-Fan Bai et al.· 1 citation
Large language models (LLMs) are high-value assets that can be derived through redeployment, fine-tuning, quantization, or further alignment. Because deployed LLMs are commonly exposed only through query APIs, ownership verification must often rely on black-box text responses. This setting is difficult: generations are...
Yu-Tong Wu, Xiao-Fan Bai, Shixin Li et al.· 1 citation
This work introduces \method, an evaluation-free compressor for complete, progressively loaded skill bundles, which leaves the agent harness unchanged and emits an ordinary directory and preserves routing, so every required file and directly callable entry remains reachable after rewriting.
Xiaofan Bai, Chao Liu, Hong-Qiang Lin et al.· 0 citations
SkillZip is presented, an evaluation-free method that compresses a skill by finding its shortest faithful structural explanation, subject to a hard coverage constraint for every extracted trigger, workflow edge, tool requirement, obligation, and output field.
Xiao-Fan Bai, Hong-Qiang Lin, Chao Liu et al.· 2 citations
SkillBoost is proposed, a three-stage framework that mitigates both risks: structured exploitation localizes observed failures to editable skill components, prior-guided exploration draws on prior knowledge in the LLM to generate diverse repair candidates, and verified acceptance commits a candidate only when it improv...
Hong-Qiang Lin, Chao Liu, Xiaofan Bai et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.