Skip to content

Author

Huan-Le Xu

We have 3 of 20 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

Poseidon: DAG-Guided Parallelism Search for LLM Pre-Training on Heterogeneous Clusters

With the rapid advancement of accelerator technologies, pre-training large language models (LLMs) on heterogeneous accelerator clusters has become increasingly crucial for maximizing hardware utilization. Existing systems, however, suffer from inaccurate training time modeling, which undermines the parallelization optimizations built upon it. Moreover, for current approaches, the vast configuration search space makes exhaustive exploration infeasible, forcing a trade-off between search time and training efficiency. To overcome these limitations, we introduce Poseidon, an efficient and scalable LLM training framework designed with heterogeneity awareness. Its core is an explicit training time model based on a directed acyclic graph. Building on this graph, Poseidon employs two efficient, theoretically grounded strategies: stage-level pruning via early stopping with partial estimation, and layer-to-stage mapping exploiting a ridge-like distribution pattern. These strategies reduce the search space without sacrificing optimal training efficiency. Experiments on heterogeneous clusters show that Poseidon improves training throughput by up to $2.76\times$ over state-of-the-art systems.

Xiao-Song Chen, Shao-Heng Nie, Zhong-Min Zhao et al. · 0 citations
Book Open access Jul 2026

Cremes: Cost-Efficient and Reliable Microservice Execution on Spot Instances

Cremes is proposed, an adaptive and cost-efficient scaling framework that ensures microservice recovery within the spot instance grace period and maintains SLO violation rates under preemptible environments below 6.7%.

Liao Chen, Chenyu Lin, Junlin Chen et al. · 0 citations
Preprint Aug 2026

psRL: Efficient Training for Agentic AI via Training-Time Prefix Sharing

This paper proposes psRL (prefix sharing for RL), a new training system for agentic AI designed to exploit prefix redundancy among training samples, and introduces two novel prefix-sharing mechanisms that enable flexible, fine-grained workload distribution across GPU workers.

Mian-Jie Yu, Zizhao Mo, Huanyu Qu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.