Verifiable instruction-following benchmarks often express each constraint through one fixed template. We test whether scores remain stable when the operational requirement is unchanged but its wording varies. We introduce WISE, a matched evaluation suite and reporting protocol instantiated on exact word count, keyword...
Qi Zhan, Seoyeon Jang, Zi-Han Dong et al.· 0 citations
Computer-use agents are increasingly evaluated in realistic desktop environments, but existing benchmarks provide limited coverage of professional engineering workflows whose outputs are persistent, structured artifacts. Mechanical computer-aided design (CAD) is a particularly demanding setting: an agent must manipulat...
Zi-Han Dong, Yuan-Zhe Liu, Zhi-Yuan Ma et al.· 0 citations
Results suggest that AAPT performs best when candidate actions can be enumerated in advance, whereas reactive execution remains stronger when they cannot, and on an external benchmark, AAPT matches the overall performance of a reactive baseline, although the two methods exhibit complementary strengths.