This paper examines how skills may help address bottlenecks of current agents and how they may expand agent capabilities through reusable domain procedures loaded at inference time and outlines open questions in skill construction, composition, evaluation, portability, governance, and security.
Hanwen Xing, Haomin Zhuang, Xuandong Zhao et al.· Proceedings of the 32nd ACM...· 7 citations· ⚡1
As Large Language Model (LLM) agents have demonstrated broad competence, but they still struggle in specialized, real-world workflows. Existing approaches such as RAG, fine-tuning and tool integration improve knowledge access, model adaptation, and external functionality, yet they do not fully address a central gap: the absence of reusable procedural knowledge for carrying out domain tasks reliably. This paper examines the emerging notion of agent skills as a possible abstraction for addressing that gap. Agent Skills are modular packages of domain-specific procedural knowledge that can be injected at inference time. Intuitively, a skill is like a cooking recipe for an agent: it does not provide new ingredients or tools, but specifies how available resources should be combined to achieve a desired outcome. A community-driven skills ecosystem is already emerging at remarkable speed, with early evidence of meaningful performance gains across multiple domains. However, their value and limits remain open questions. We examine how skills may help address bottlenecks of current agents and how they may expand agent capabilities through reusable domain procedures loaded at inference time. We then outline open questions in skill construction, composition, evaluation, portability, governance, and security, and conclude with a call for contribution. Our goal is not to present skills as a settled solution, but to clarify their promise, limits, and the questions that must be answered before they can become a principled foundation for future agent systems.
Hanwen Xing, Haomin Zhuang, Xuandong Zhao et al.· Proceedings of the 32nd ACM...· 6 citations· ⚡1
MatrAIx is introduced, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users and provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.
Xiaomin Li, Yuexing Hao, Jian Hou et al.· 1 citation
This work introduces a novel approach, \textit{CanaryTrace}, to safeguard the ownership of text datasets and effectively detect unauthorized use by RA-LLMs, and demonstrates high query efficiency, detectability, and consistency, along with minimal perturbation to the original dataset, all without compromising the performance of the RAG system.
Yepeng Liu, Xuandong Zhao, D. Song et al.· arXiv.org· 15 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.