Workspace agents combine large language models with execution harnesses to perform stateful, multi-step tasks that access or modify external resources. Existing benchmarks leave gaps in executable coverage of their runtime security risks, while evolving model capabilities, harnesses, tools, and threats motivate benchma...
Shi Kuang, Xue-Mei Luo, Kun Liu et al.· 0 citations
A Mixture-of-Experts framework that combines multiple policies, leveraging their complementary strengths to form a more robust exploration policy is proposed, and the evaluation on the Air Traffic benchmark shows that this proposal significantly increases the number of solvable instances.
Toshihide Ubukata, Mingyue Zhang, Zhi-Yao Wang et al.· SEAMS@ICSE· 1 citation
TeCoR-UAV achieves better bi-objective trade-offs in most medium- and large-scale scenarios, as well as in topologically constrained scenarios, and improves service quality by an average of 18.5 percentage points, indicating its scenario adaptability and potential for practical application.
Buyang Ding, Weijun Ni, Yixing Luo et al.· Electronics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.