Skip to content

Author

Zhiheng Xi

We have 10 of 85 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

An AI System for Autonomous Algorithm Evolution in Drug Development

Artificial intelligence (AI) is increasingly permeating the drug development pipeline. Numerous algorithms for accelerating this multi-stage and multi-task process have been constructed, which depends heavily on expert design and labor-intensive task-specific optimization. Given that AI-driven acceleration of drug deve...

Zhi-Meng Zhou, Yang Nan, Min-Jie Mou et al. · 0 citations
Preprint Aug 2026

A Token-Level Analysis of Sampled-Token Reverse-KL On-Policy Distillation

On-policy distillation (OPD) supervises a student on its own trajectories with token-level signals from a frozen teacher, yet how a sampled loss allocates updates across tokens remains poorly understood. We analyze the gradient of the per-token K2 estimator of reverse KL with respect to the student logits. The $\ell_1$...

Bing Shao, Jia-Zheng Zhang, Long Ma et al. · 5 citations
#artificial intelligence Review Sep 2026

Atria Dawn: The Dawn of Agentic Superintelligence

Atria Dawn Preview is introduced, a foundation agentic language model designed for scientific research and engineering workflows, with the goal of expanding the frontier of agent productivity in the real world.

Honglin Guo, Tao Gui, Kun Cai et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents

Sci-MMR is introduced, a benchmark for multi-step evidence-grounded scientific reasoning built on structured argument graphs linking scientific claims, citation-grounded knowledge, visual evidence, and supporting regions, and it is found that current answer-centric benchmarks substantially overestimate the evidence-gro...

Jia-Qiang Li, Ya-Jie Yang, Zhi-Heng Xi et al. · 0 citations
Conference Open access Sep 2026

MathCritique: Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision

A critique-in-the-loop self-improvement method that incorporates critique-based supervision into the actor’s self-training process and improves the actor’s exploration efficiency and solution diversity, especially on challenging queries, leading to a stronger actor model.

Zhi-Heng Xi, Dingwen Yang, Jixuan Huang et al. · 0 citations

Prefix-Adaptive Block Diffusion for Efficient Document Recognition

The Prefix-Adaptive Block Diffusion Model (PA-BDM) is proposed, which replaces intra-block bidirectional denoising with causal denoising from prefix to suffix and treats the block size as a maximum candidate range rather than a fixed commitment unit.

Ming-Xu Chai, Zi-Yu Shen, Chen-Yu Liu et al. · 0 citations

JFTA-Bench: Evaluate LLM's Ability of Tracking and Analyzing Malfunctions Using Fault Trees

A novel textual representation of fault trees is proposed, and a benchmark for multi-turn dialogue systems that emphasizes robust interaction in complex environments is constructed, evaluating a model's ability to assist in malfunction localization.

Yuhui Wang, Zhi-Xiong Yang, Ming Zhang et al. · 0 citations
Preprint Aug 2026

CAFE: Self-Improving Search Agents Need Co-Evolving Feedback

CAFE (Coupled Agent--Feedback Evolution), a framework in which a shared-parameter model alternates between search-agent and critic roles, is introduced, suggesting that a self-improving search agent needs feedback that co-evolves with the policy it guides.

Bo-Yang Liu, Senjie Jin, Pei-Xin Wang et al. · 1 citation
Conference Open access 2026

Counteracting the Matthew Effect in Self-Improvement of LVLMs through Head-Tail Re-balancing

To mitigate a critical imbalance during the exploration-and-learning process, this work approaches head-tail re-balance during the exploration-and-learning process from two perspectives: distribution-reshaping and trajectory-resampling.

Xin Guo, Zhiheng Xi, Yiwen Ding et al. · 1 citation
Conference Open access Jul 2026

AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments

AgentGym2 is presented, a new evaluation framework with task instances grounded in real-world end-to-end working demands that measures agents'ability to execute end-to-end procedures, discover tools via exploration, compose tools for unseen tasks, and remain robust to noisy and underspecified information.

Zhiheng Xi, Dingwen Yang, Jiaqi Liu et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.