SwarmBench is proposed, a benchmark that evaluates model performance from multiple perspectives, including accuracy, efficiency, cost, and process quality, and SwarmExp is proposed, a simple yet effective method based on experience extraction and experience replay, which consistently improves the orchestration performance of large language models.
Jin Gao, Zhuoran Jin, Tianyi Men et al.· 0 citations
RuleWeaver is introduced, a benchmark construction framework for evaluating rule-centered scenario reasoning that starts from corpus-derived IF-THEN Meta Rules, progressively augments them into complex rules, and composes these rules into rule-centered scenario QA instances.
Bohan Yu, Shi-Yang Li, Pengfei Cao et al.· 0 citations
DynaRule is proposed, an end-to-end framework that injects the given rules into the KV cache and turns retrieval into an internal, learnable, step-wise process, and can re-attend to the most relevant rules at each step, dynamically replacing outdated ones to support more stable multi-step reasoning.
Bohan Yu, Pengfei Cao, Chen Han et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.