Skip to content

Author

Binhang Yuan

We have 7 of 25 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Oct 2026

Coda: Exploiting Admission Flexibility for Coding-Agent Serving

Coding agents powered by large language models (LLMs) repeatedly alternate between model inference and tool calls, creating long-lived sessions with reusable key-value (KV) states and asynchronous request resumptions. Logical readiness, however, does not ensure efficient admission in a shared serving system. Through di...

You-He Jiang, Fang-Cheng Fu, Bin-Hang Yuan et al. · 0 citations
Preprint Oct 2026

ePACT: Energy-Performance-Aware Commitment Tracking for LLM Serving

Reducing LLM serving energy does not by itself guarantee lower deployment cost when electricity procurement exposes operators to unfavorable deviations from preset commitments. We study hourly commitments with positive, potentially asymmetric costs for overuse and underuse, and formulate energy-Performance-Aware Commit...

You Peng, You-He Jiang, Chen Wang et al. · 0 citations
#artificial intelligence Preprint Oct 2026

D2K-Bench: Can LLM Agents Turn Expert Designs into Efficient GPU Kernels?

GPU kernels generated by large language model (LLM) agents can remain less efficient than expert implementations, but runtime alone does not reveal how the gap relates to design discovery and implementation. We introduce D2K-Bench, a diagnostic benchmark of 26 tasks and 85 workloads that measures how effectively agents...

Dai-Feng Li, Huiqiang Jiang, Chengruidong Zhang et al. · 0 citations
#machine learning Preprint Sep 2026

Denoising Surface: Modeling and Predicting Inference Cost for Diffusion LLM Serving

As diffusion large language models (dLLMs) become more capable, they are moving from research settings to real-world \textit{serving}, where request management (such as scheduling and resource allocation) relies on accurate estimation of per-request inference cost. However, common cost proxies fall short for dLLMs: out...

Hao-Yu Zheng, Fang-Cheng Fu, Bin-Hang Yuan et al. · 0 citations
Preprint Sep 2026

AReaL-TIK: Stateful Agentic Optimization of Unified RL Kernels through an Optimization IR

Reinforcement learning (RL) post-training often uses distinct GPU kernels for rollout and policy update. In synchronous PPO and GRPO, numerical disagreement can perturb ratios between current token probabilities and those assigned during rollout. Recomputing rollout log-probabilities with the policy-update backend avoi...

Ran Yan, You-He Jiang, Jia-Yi Nie et al. · 0 citations
Preprint Aug 2026

AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement Learning

AReaL-DTE is presented, a snapshot-free Delta Transfer Engine that translates inference-visible weight sparsity into end-to-end system efficiency and achieves speedups of up to 19.9x over ByteCheckpoint and 3.2x over PULSE across clusters, and up to 7.6x and 7.4x within a cluster.

Yingqi Peng, Jia-Wei Zhang, Wenhao Zhou et al. · 1 citation

OpenTela: Unifying Decentralized Computing Resources for Heterogeneous LLM Serving (Operational Systems)

OpenTela is presented, a user-space orchestration overlay that turns existing fragmented HPC clusters into a unified, cross-institutional serving platform and provides a replicable blueprint for other sovereign AI initiatives to harness their own federated GPU infrastructure.

Xiaozhe Yao, Youhe Jiang, Ilia Badanin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.