GraphForge, an evidence-graph based framework that grounds both the task and its verification in real files, and rejects fine-tuning on the SFT model's own rollouts yields further improvements on all three benchmarks, suggesting that the rubrics provide a useful selection signal.
Qi-Sheng Su, Han-Chen Wang, Guang Zhu et al.· 0 citations
Recent efforts to scale tool-use post-training have largely centered on the synthesis of executable environments, which constitute only one component of a broader agentic interaction system comprising the environment, task, agent harness, and evaluator. Scaling environments in isolation, however, does not guarantee com...
Bo Mao, Hang He, Lin-Ting Wang et al.· 0 citations
EvoIn is an agent fine-tuning framework that bridges evolution and internalization, and consistently enables agents to learn stronger decision-making procedures, raising the pass rate by 10.9 points in-domain and by 9.2 points out-of-domain.
Shi-Han Dou, Shao-Hua Liu, Zhong-Hang Lu et al.· 0 citations
ExplorationBench turns the wicked problem of evaluating scientific exploration into a concrete and tractable framework built on verifiable Alien Worlds, and finds that the strongest systems can acquire and apply unfamiliar rules, while performance varies substantially across trajectories and continued exploration can s...
Ming Zhang, Zhen Xiang, Pei-Zhong Gao et al.· 0 citations
Producing task-specific large language models requires discovering effective training strategies through experimentation. Automated fine-tuning systems have made this experimentation feasible with far less manual effort. However, these systems are stateless: each search discards its discovered strategies, dataset insig...
Hao-Ran Zhao, Wei Du, Dingwen Yang et al.· 0 citations
On-policy distillation (OPD) supervises a student on its own trajectories with token-level signals from a frozen teacher, yet how a sampled loss allocates updates across tokens remains poorly understood. We analyze the gradient of the per-token K2 estimator of reverse KL with respect to the student logits. The $\ell_1$...
Bing Shao, Jia-Zheng Zhang, Long Ma et al.· 5 citations
Experiments on document parsing benchmarks show that SCVER improves robustness under reduced input resolution and achieves a better accuracy-efficiency trade-off, demonstrating the effectiveness of on-demand visual evidence retrieval for fine-grained perception.
Ming-Xu Chai, Chen-Yu Liu, Zi-Yu Shen et al.· 0 citations
Atria Dawn Preview is introduced, a foundation agentic language model designed for scientific research and engineering workflows, with the goal of expanding the frontier of agent productivity in the real world.
MuseCritic is introduced, a semi-scalar reward model that generates a natural-language critique covering five aesthetic dimensions and uses it as an intermediate representation to predict continuous reward scores, showing that critique-conditioned reward modeling reduces scoring error and provides an effective optimiza...
Jiabao Zhuang, Changhao Jiang, Hanchen Wang et al.· 0 citations
Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scattered across multiple applications (apps) to complete user instructions. However, due to the lack of dedicated benchmarks, their capabilities remain poorly understood. To address thi...
Junjie Ye, Zhuohui Sheng, Shao-Hua Liu et al.· 0 citations
Sci-MMR is introduced, a benchmark for multi-step evidence-grounded scientific reasoning built on structured argument graphs linking scientific claims, citation-grounded knowledge, visual evidence, and supporting regions, and it is found that current answer-centric benchmarks substantially overestimate the evidence-gro...
Jia-Qiang Li, Ya-Jie Yang, Zhi-Heng Xi et al.· 0 citations
NovGauge is presented, a human-anchored benchmark for fine-grained novelty assessment diagnosis, and a cascading diagnostic pipeline that verifies per-dimension correctness, evidence grounding, and logical support is proposed.
Guo-Qiang Zhang, Ke-Xin Tan, Ming Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.