EvoPathBench establishes capability-level process evaluation as a foundation for analyzing self-evolution, identifying candidate evaluation and selection as key targets for improvement.
Hong-Qiang Lin, Chao Liu, Xiao-Fan Bai et al.· 1 citation
This work introduces \method, an evaluation-free compressor for complete, progressively loaded skill bundles, which leaves the agent harness unchanged and emits an ordinary directory and preserves routing, so every required file and directly callable entry remains reachable after rewriting.
Xiaofan Bai, Chao Liu, Hong-Qiang Lin et al.· 0 citations
SkillZip is presented, an evaluation-free method that compresses a skill by finding its shortest faithful structural explanation, subject to a hard coverage constraint for every extracted trigger, workflow edge, tool requirement, obligation, and output field.
Xiao-Fan Bai, Hong-Qiang Lin, Chao Liu et al.· 2 citations
As large language models are increasingly deployed in real-world systems, safety failures can still lead to harmful outputs and dangerous misuse. We argue that the essence of safety is adversarial: many failures arise not from natural inputs alone, but from strategic attempts to evade model policies and safeguards. How...
Ting Ma, Xiufeng Huang, Benlei Cui et al.· 0 citations
SkillBoost is proposed, a three-stage framework that mitigates both risks: structured exploitation localizes observed failures to editable skill components, prior-guided exploration draws on prior knowledge in the LLM to generate diverse repair candidates, and verified acceptance commits a candidate only when it improv...
Hong-Qiang Lin, Chao Liu, Xiaofan Bai et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.