SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe
This work formalizes skill optimization via Zeroth-Order (ZO) optimization, mapping classical counterparts (central difference, trust regions) to recent literature to establish three principles for convergence and generalization: file-system-based trajectory exploration, consensus attribute mining, and independent validation gating.