Coding agents powered by large language models (LLMs) repeatedly alternate between model inference and tool calls, creating long-lived sessions with reusable key-value (KV) states and asynchronous request resumptions. Logical readiness, however, does not ensure efficient admission in a shared serving system. Through di...
You-He Jiang, Fang-Cheng Fu, Bin-Hang Yuan et al.· 0 citations
Reducing LLM serving energy does not by itself guarantee lower deployment cost when electricity procurement exposes operators to unfavorable deviations from preset commitments. We study hourly commitments with positive, potentially asymmetric costs for overuse and underuse, and formulate energy-Performance-Aware Commit...
You Peng, You-He Jiang, Chen Wang et al.· 0 citations
Reinforcement learning (RL) post-training often uses distinct GPU kernels for rollout and policy update. In synchronous PPO and GRPO, numerical disagreement can perturb ratios between current token probabilities and those assigned during rollout. Recomputing rollout log-probabilities with the policy-update backend avoi...
Ran Yan, You-He Jiang, Jia-Yi Nie et al.· 0 citations
OpenTela is presented, a user-space orchestration overlay that turns existing fragmented HPC clusters into a unified, cross-institutional serving platform and provides a replicable blueprint for other sovereign AI initiatives to harness their own federated GPU infrastructure.
Xiaozhe Yao, Youhe Jiang, Ilia Badanin et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.