GPU kernels generated by large language model (LLM) agents can remain less efficient than expert implementations, but runtime alone does not reveal how the gap relates to design discovery and implementation. We introduce D2K-Bench, a diagnostic benchmark of 26 tasks and 85 workloads that measures how effectively agents...
Dai-Feng Li, Huiqiang Jiang, Chengruidong Zhang et al.· 0 citations
Language-model agents increasingly face long-horizon tasks with evolving state, interdependent decisions, and delayed outcomes. Scaling their training requires diverse agentic environments, dependable outcome signals, and low extension cost. Existing generation pipelines commonly construct an environment before definin...
Xin-Jie Shen, Wei Fan, Xu-Dong Guo et al.· 0 citations
This work introduces RecreationWorld, a five-platform framework built around recreation: given a running reference, an agent must discover its behavior and build a faithful implementation with no prescribed workflow, and introduces RecreationBench, comprising 250 diverse tasks across domains and platforms.
Shuai Bai, Jia-Yong Deng, Si-Cheng Fan et al.· 1 citation
The proposed MIMIC framework fundamentally transforms algorithms into verifiable reasoning trajectories through narrative fusion, code-guided test synthesis, and dynamic code instrumentation, demonstrating that the procedural rigor of executable code can effectively unlock and enhance the generalized reasoning capabili...
Jin-Yang Zhang, Wei-Bin Liao, Ke-Qin Bao et al.· 0 citations
We present Qwen-Drive-1.0, an initial step towards a vision-language foundation model for autonomous driving. Qwen-Drive-1.0 retains the architecture of the pretrained vision-language model (VLM) and integrates 3D perception, visual question answering, and motion planning within a unified framework. An external bird's-...
Xin Zhou, Zongchuang Zhao, Zhibo Yang et al.· 2 citations
Aligning over 210k synthetic preference samples, BridgeAlign enables Qwen3-8B to achieve the best average across 17 benchmarks against 11 strong baselines, leading on both human-preference and knowledge-based capabilities at once, with no trade-off between them.
Terminal-Universe is introduced, a framework which turns each trajectory into a reusable environment and explores it for synthesizing new tasks and continued interactions, and produces 37.3k task-sufficient environments.
Jie Wu, Zhen-Ru Zhang, Bei-Chen Zhang et al.· 3 citations
COPD is proposed, a contrastive OPD framework that substantially reduces reasoning length without compromising model performance and consistently improves efficiency across different tasks and model scales.
Jiacheng Ruan, Jun Tang, Wenzhen Yuan et al.· arXiv.org· 2 citations
It is demonstrated that the hidden states of probed answers more effectively differentiate distinct solution paths than semantic embeddings, and the perplexity of probed answers serves as a practical proxy for reasoning correctness.
Yi Fang, Quek Shen, Chengping Li et al.· 0 citations
FinIndices is a large-scale benchmark evaluating data-processing fidelity over uncropped financial statements (up to 32K tokens) and yields substantial zero-hint gains, validating that structured logic can be partially restored via data-centric alignment.
Quantile-Guided Density Estimation (QGDE), which approximates this distribution with multiple quantile trends and uses local density weighting to produce token-level estimates and suggests that released tokenizer vocabularies provide a useful signal for fine-grained corpus estimation beyond coarse composition inference...
Qingjie Zhang, Xing-Zhang Ren, Zi-Xuan Chen et al.· 0 citations
The architecture and the Muon optimizer together shift the optimal learning rate and batch size upwards, render batch-size warmup unnecessary, and substantially improve stability under stress tests.
Zi-Han Qiu, Ze-Kun Wang, Xiao Li et al.· 24 citations· ⚡3
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.