Expert-written natural language skills can improve tool-using agents, yet agent-authored skills perform 8-11 points worse than using no skill. This gap suggests that following procedural guidance and improving it from execution evidence are distinct capabilities. Inference time loops can repair skills but do not improv...
Kang Peng, Zhi-Wei Zhang, Yichen Zhang et al.· 1 citation
Large language model post-training generates self-generated rollouts through reinforcement learning and on-policy distillation, yet this experience is often treated as stale once the policy advances. Historical rollouts can remain compatible with a later policy while preserving behaviors that the policy no longer expre...
Zhi-Wei Zhang, Hua-Yu Deng, Fei Zhao et al.· 0 citations
MatrAIx is introduced, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users and provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.
Xiaomin Li, Yuexing Hao, Jian Hou et al.· 1 citation
This work introduces PolyWorkBench, a benchmark for evaluating LLM agents on multilingual long-horizon workplace workflows, and proposes a hybrid framework that combines structural grading, executable verification, and LLM-based semantic assessment to enable comprehensive evaluation.
Benchmark evaluations reveal that agent performance varies substantially across languages and drops sharply on the harder cross-lingual tasks, and analysis shows that multilingual execution exposes systematic failure modes across planning, tool interaction, and decision-making in long-horizon agents.
Hongliang Li, Yijin Liu, Zhiwei Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.