Skip to content

Author

Dayiheng Liu

We have 17 of 114 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Oct 2026

D2K-Bench: Can LLM Agents Turn Expert Designs into Efficient GPU Kernels?

GPU kernels generated by large language model (LLM) agents can remain less efficient than expert implementations, but runtime alone does not reveal how the gap relates to design discovery and implementation. We introduce D2K-Bench, a diagnostic benchmark of 26 tasks and 85 workloads that measures how effectively agents...

Dai-Feng Li, Huiqiang Jiang, Chengruidong Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms

Language-model agents increasingly face long-horizon tasks with evolving state, interdependent decisions, and delayed outcomes. Scaling their training requires diverse agentic environments, dependable outcome signals, and low extension cost. Existing generation pipelines commonly construct an environment before definin...

Xin-Jie Shen, Wei Fan, Xu-Dong Guo et al. · 0 citations

RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents

This work introduces RecreationWorld, a five-platform framework built around recreation: given a running reference, an agent must discover its behavior and build a faithful implementation with no prescribed workflow, and introduces RecreationBench, comprising 250 diverse tasks across domains and platforms.

Shuai Bai, Jia-Yong Deng, Si-Cheng Fan et al. · 1 citation
#artificial intelligence Preprint Sep 2026

The Imitation Game: When LLMs Learn to Reason Like Programs via Code-Centric Reasoning Data Synthesis

The proposed MIMIC framework fundamentally transforms algorithms into verifiable reasoning trajectories through narrative fusion, code-guided test synthesis, and dynamic code instrumentation, demonstrating that the procedural rigor of executable code can effectively unlock and enhance the generalized reasoning capabili...

Jin-Yang Zhang, Wei-Bin Liao, Ke-Qin Bao et al. · 0 citations
Preprint Aug 2026

Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

We present Qwen-Drive-1.0, an initial step towards a vision-language foundation model for autonomous driving. Qwen-Drive-1.0 retains the architecture of the pretrained vision-language model (VLM) and integrates 3D perception, visual question answering, and motion planning within a unified framework. An external bird's-...

Xin Zhou, Zongchuang Zhao, Zhibo Yang et al. · 2 citations
Jul 2026

BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences

Aligning over 210k synthetic preference samples, BridgeAlign enables Qwen3-8B to achieve the best average across 17 benchmarks against 11 strong baselines, leading on both human-preference and knowledge-based capabilities at once, with no trade-off between them.

Ru Peng, Haokai Xu, Xijun Gu et al. · 0 citations
Jul 2026

Contrastive On-Policy Distillation

COPD is proposed, a contrastive OPD framework that substantially reduces reasoning length without compromising model performance and consistently improves efficiency across different tasks and model scales.

Jiacheng Ruan, Jun Tang, Wenzhen Yuan et al. · 2 citations
Jul 2026

Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements

FinIndices is a large-scale benchmark evaluating data-processing fidelity over uncropped financial statements (up to 32K tokens) and yields substantial zero-hint gains, validating that structured logic can be partially restored via data-centric alignment.

Xin Tong, Xuanming Zhang, Tian-Yi Tang et al. · 0 citations
#natural language process... Preprint Aug 2026

Can Released LLM Vocabularies Support Token-Level Estimation of Hidden Corpora?

Quantile-Guided Density Estimation (QGDE), which approximates this distribution with multiple quantile trends and uses local density weighting to produce token-level estimates and suggests that released tokenizer vocabularies provide a useful signal for fine-grained corpus estimation beyond coarse composition inference...

Qingjie Zhang, Xing-Zhang Ren, Zi-Xuan Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.