E-Commerce Bench is introduced, the first open-source benchmark that integrates multi-round counterpart negotiation and dynamic events into a year-long business operation, and it is found that no single model dominates.
Wei Fan, Xin-Jie Shen, Xu-Dong Guo et al.· 2 citations
WILDTRACE is introduced, a benchmark of 481 tasks over 214 naturally occurring long-form sources such as technical incident reports and lesser-known literary narratives, where all evidence trails arise from the document's own causal, temporal, and narrative logic.
Zixin Chen, Peng Liu, Haobo Li et al.· 0 citations
H-Scale is a lightweight post-processing method for NVFP4 per-group scale refinement that selects hardware-valid group scales using a diagonal second-order proxy derived from calibration activations, thereby targeting layer output perturbation more directly.
Hao Yu, Zheng Li, Dayiheng Liu et al.· 1 citation· ⚡1
Qwen-CUA is introduced, a native computer-use agent with a 397B-A17B Qwen mixture-of-experts backbone that outperforms Qwen3.7 and remains competitive with leading proprietary systems, and scalable verifiable interaction and hybrid tool use as key directions.
Dunjie Lu, Shuai Bai, Tianyi Bai et al.· 2 citations
These results show that current agents are still far from professional-level computer use: rather than stumbling on basic GUI control or coding, they lose track of constraints, miss information that arrives mid-task, guess rather than ask the user, and skip verification, struggling most when a task hinges on hidden sta...