Sep 2026· Proceedings of the International Conference on Parallel Processing· 0 citations· 40 references
TL;DR
OptPipe is presented, a unified framework that jointly optimises partitioning and scheduling for pipeline parallelism and introduces a memory-aware directed acyclic graph (DAG) that captures both task dependencies and the lifetime of intermediate tensors, enabling explicit reasoning about the trade-off between execution efficiency and memory usage.
Abstract
Pipeline parallelism (PP) is widely adopted for distributed training of neural networks, but its efficiency is limited by pipeline bubbles that leave devices idle. A variety of schedules have been proposed to reduce these bubbles, yet most are hand-crafted, rely on idealised timing assumptions as well as rigid structural constraints, and cannot adjust to memory limitations. Moreover, stage partitioning, task scheduling, and memory management are often optimised in isolation, despite being tightly interdependent. In this paper, we present OptPipe, a unified framework that jointly optimises partitioning and scheduling for pipeline parallelism. We introduce a memory-aware directed acyclic graph (DAG) that captures both task dependencies and the lifetime of intermediate tensors, enabling explicit reasoning about the trade-off between execution efficiency and memory usage. Building on this representation, we present a mathematical formulation for pipeline parallelism that simultaneously determines stage partitioning, device assignment, task ordering, and memory feasibility. Unlike prior methods, our approach imposes no structural constraints on pipeline configurations and adapts naturally to a wide range of memory budgets and workload characteristics. Experiments on GPT-3 and LLaMA-3.2 show that OptPipe improves training throughput by up to 72.5%, 37.8%, and 38.0% over 1F1B, Chimera, and ZB-V, while remaining feasible under memory budgets where these baselines fail.
OmniPipe is proposed, a flexible bidirectional multi-pipeline parallelism scheme for unified dense and MoE LLM training that minimizes the pipeline bubble ratio while effectively overlapping EP communication with computation, enabled by the flexible and scalable parallelism scheme of bidirectional pipelines.
Jun Li, Zhi Ma, Shi-Gang Li· Proceedings of the Internati...· 0 citations
Large Language Model (LLM) inference has become the dominant workload in modern AI systems, requiring serving infrastructures to maximize throughput while meeting strict latency Service-Level Objectives (SLOs). Since state-of-the-art LLMs exceed the compute and memory capacity of a single GPU, inference is commonly dis...
HDA-MoE is presented, a framework that optimizes MoE execution on NMP architectures through hybrid parallel deployment and runtime scheduling and integrates an offline hybrid parallel mapping algorithm with an online dynamic and adaptive scheduling mechanism to reduce communication overhead while improving computation...
Hao-Chen Huang, Shu-Zhang Zhong, Sheng-Xuan Qiu et al.· IEEE Transactions on Compute...· 0 citations
Poseidon is an efficient and scalable LLM training framework designed with heterogeneity awareness, which employs two efficient, theoretically grounded strategies: stage-level pruning via early stopping with partial estimation, and layer-to-stage mapping exploiting a ridge-like distribution pattern.
Xiao-Song Chen, Shao Nie, Zhong-Min Zhao et al.· 0 citations
EPIC mitigates imbalance via performance-aware expert migration and runtime expert activation, and then improves communication with topology-adaptive transport kernels and fine-grained computation-communication overlap.
Jia-Min Cao, Qingxu Li, Yaozhong Liu et al.· Conference on Applications,...· 2 citations· ⚡1
The past decade has seen a variety of novel computational accelerators and disruptive memory systems, providing, e.g., in-memory or tensor processing capabilities. Executing tasks and placing data on these can drastically improve application performance, but only when done right - otherwise, performance may also degrad...
Marcel Lütke Dreimann, B. Friesel, Lilith Lenze et al.· Proceedings of the 4th Works...· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduJul 15, 2026
Assistant Professor Pat Pataranutaporn describes a new interface that lets everyday users glimpse inside an AI's neural network before their chatbot ever says a word.
Microsoft Research Blog· microsoft.comJul 13, 2026
Cryptographic code supports vital protections in modern computing systems. Learn how a new method helps verify code as developers write it while preserving speed and adaptability as it gets implemented and evolves. The post Verifying Rust cryptography in SymCrypt, from standards to code appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduJul 6, 2026
PhD student Rachel Sava, winner of the Envisioning the Future of Computing Prize, explores transformative improvements and dystopian risks of neural technology.