Rubric-based evaluation is widely used to assess LLM-based systems by decomposing response quality into task-specific scoring criteria. However, automatically generating rubrics that reliably capture task-specific quality requirements remains challenging. We introduce Mubric, a mutation testing-guided approach to rubri...
Jiayuxuan Yang, Jie M. Zhang, Yiling Lou et al.· 0 citations
A reuse paradox is revealed: although skills are intended to be easily imported and reused, developers spend a lot of effort rewriting what the skills do, fixing skill discoverability, and translating them for different tools and languages, indicating a need for better abstractions, standardized interfaces, and automat...
Xinjian Wu, Jingzhi Gong, Gunel Jahangirova et al.· 0 citations
This paper proposes SafeKeep, an inference-time safeguard that decouples safety judgment from tool execution: it assesses requests using flattened textual tool specifications while retaining the original schema-formatted specifications for execution.