Skip to content

Author

Sungjae Lee

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Book Open access Jul 2026

HAT-MPI: Hierarchical Auto Tuning of MPI Inter-Node Communication on InfiniBand Clusters

Optimizing point-to-point and collective communication in HPC systems substantially impacts the performance of large-scale applications. Middleware libraries such as Message Passing Interface (MPI) rely on lower-level communication libraries, including UCX and OFI Libfabric, to enable these optimizations. However, because optimal configurations depend on the characteristics of each HPC system, users must understand system-specific performance behavior and manually tune communication parameters. In this work, we present HAT-MPI, an ML-based framework that identifies the UCX Eager-Rendezvous protocol threshold and the optimal collective algorithm on InfiniBand HPC clusters. Using hardware characteristics and performance data collected across clusters at different scales, HAT-MPI predicts the UCX protocol threshold and the optimal Alltoall/Allreduce algorithm for previously unseen InfiniBand clusters. This paper also introduces two complementary techniques that improve the conventional data collection and evaluation workflow, namely LLM-assisted HPC hardware detection and Top-5% candidate algorithm evaluation. Experimental results show that the predicted UCX protocol threshold achieves an average of 31% performance improvement in regions where a gap exists between the predicted and default thresholds, while the algorithm selection attains 84.9% accuracy with Top-5% evaluation, demonstrating that HAT-MPI delivers highly accurate predictions.

Sungjae Lee, Sam Tilford, D. Panda · 0 citations
Book Open access Aug 2026

Toward WAN-Aware LLM Training Across Heterogeneous, Geo-Distributed Sites

Large Language Model (LLM) training is increasingly concentrated in homogeneous datacenters, while private data and underutilized GPUs across universities, laboratories, and edge sites remain difficult to use. This extended abstract presents preliminary results from a geo-distributed LLM training prototype that treats networking constraints as first-order design concerns. The prototype connects three heterogeneous GPU sites via cloud-hosted parameter servers, outbound-only gRPC streams, two-stage delta compression (INT8 quantization + Huffman coding, achieving up to 4× payload reduction), and fault-tolerant rejoin. In real deployments, GPT-2 Medium pretraining achieves stable loss reduction and reaches the target loss 15.2% faster in wall-clock time than the best tested baseline; Llama3-1B pretraining remains stable under larger communication pressure; and cross-site latency traces reveal site-dependent WAN spikes of up to 200s. These results motivate adaptive networking support for synchronization, compression, placement, telemetry, and recovery in geo-distributed LLM training.

Ziyue Luo, Jiaxuan Cai, Cedric Le Denmat et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.