Large language model (LLM) agents can benefit from reusable skills distilled from prior task experience, yet existing skill optimization methods often rely on costly execution-based evaluation and substantial task data. We introduce \textbf{COBRA-Skills}, an efficient framework that formulates skill optimization as budgeted sequential optimization over a dynamically evolving candidate space. COBRA-Skills couples contextual-bandit-guided prioritization with evidence-grounded skill evolution, selectively allocating evaluations to promising or informative candidates while continually refining the skill population from execution feedback. Across six heterogeneous agent benchmarks and three target models, COBRA-Skills consistently achieves the strongest average performance among compared methods, while reducing optimization cost by 55--58\% relative to SkillOpt and using only 50 unique optimization examples per benchmark. Further analyses show that COBRA-Skills remains robust to changes in the agent harness and performs effectively when the target model itself is used for skill generation and refinement.
Ping-Chen Lu, Xiang-Yi Wang, Xiang Li et al.· 0 citations
This paper proposes algorithms that model online fair division as a contextual bandit problem and achieve provable sublinear regret and proposes algorithms that model utility is an unknown function of item-agent features.
A. Verma, Indrajit Saha, Makoto Yokoo et al.· 4 citations
This work describes the near-optimal region, the set of allocations within a specified tolerance of peak performance, which is wide even for small tolerances, widens with model scale, and transfers reliably from small proxy models to large target models.
Jingtan Wang, A. Verma, Xiaoqiang Lin et al.· 0 citations
This work introduces Power-Law Entropy Search (PLES), a computational cost-aware acquisition function built on multi-fidelity Bayesian optimization that efficiently estimates optimal hyperparameter scaling laws through adaptive experimentation using less than one-tenth of the computational budget required by conventional grid search and other baselines.
Zhiliang Chen, S. Ament, David Eriksson et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.