Skip to content

Author

Dayiheng Liu

We have 17 of 114 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Jul 2026

WildTrace: Benchmarking Natural Evidence Trails in Long-Context Reasoning

WILDTRACE is introduced, a benchmark of 481 tasks over 214 naturally occurring long-form sources such as technical incident reports and lesser-known literary narratives, where all evidence trails arise from the document's own causal, temporal, and narrative logic.

Zixin Chen, Peng Liu, Haobo Li et al. · 0 citations
#natural language process... Preprint Aug 2026

H-Scale: Hessian-Guided Scale Refinement for NVFP4 Sub-Byte LLM Inference

H-Scale is a lightweight post-processing method for NVFP4 per-group scale refinement that selects hardware-valid group scales using a diagonal second-order proxy derived from calibration activations, thereby targeting layer output perturbation more directly.

Hao Yu, Zheng Li, Dayiheng Liu et al. · 1 citation · ⚡1
Preprint Aug 2026

Qwen-CUA: Native Computer Use for (almost) Everything

Qwen-CUA is introduced, a native computer-use agent with a 397B-A17B Qwen mixture-of-experts backbone that outperforms Qwen3.7 and remains competitive with leading proprietary systems, and scalable verifiable interaction and hybrid tool use as key directions.

Dunjie Lu, Shuai Bai, Tianyi Bai et al. · 2 citations
Jun 2026

OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks

These results show that current agents are still far from professional-level computer use: rather than stumbling on basic GUI control or coding, they lose track of constraints, miss information that arrives mid-task, guess rather than ask the user, and skip verification, struggling most when a task hinges on hidden sta...

Mengqi Yuan, Zilong Zhou, Xinzhuang Xiong et al. · 11 citations · ⚡5

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.