Sep 2026· Proceedings of the 4th Workshop on Disruptive Memory Systems· 0 citations· 14 references
Abstract
The past decade has seen a variety of novel computational accelerators and disruptive memory systems, providing, e.g., in-memory or tensor processing capabilities. Executing tasks and placing data on these can drastically improve application performance, but only when done right - otherwise, performance may also degrade. Even when leaving the algorithmic part aside, this is challenging: data layout and task structure affect optimal placement, but placement decisions also affect the required data layout and task structure. Crucially, today's de facto standard of expressing high-performance computing workloads as directed acyclic graphs (DAGs) assumes a static application structure, and is therefore unsuitable for such devices. This paper proposes hDAGs, a workload definition method that is capable of expressing accelerator-specific, placement-dependent sub-structures within applications, thus addressing the aforementioned challenge. We show that hDAG-based placement can reduce application latency by up to 38 % compared to existing approaches, while also improving simulation-based latency prediction by up 61%. The performance of real-world applications can be improved by a factor of 1.20x even with simple scheduler adjustments. At the same time, hDAGs are backwards-compatible with existing scheduling strategies, and thus do not require placement algorithms to be re-designed from scratch.
OptPipe is presented, a unified framework that jointly optimises partitioning and scheduling for pipeline parallelism and introduces a memory-aware directed acyclic graph (DAG) that captures both task dependencies and the lifetime of intermediate tensors, enabling explicit reasoning about the trade-off between executio...
Ning Wang, A. Raith, Oliver Sinnen· Proceedings of the Internati...· 0 citations
PFM (PIM-as-Flexible-Memory), a dual-view memory system that decouples physical data layout from accessor-visible logical views, is presented, demonstrating its effectiveness and broad applicability as a unified memory management solution for NPU-PIM systems.
Shixin Zhao, Lian Liu, Tian Han et al.· 0 citations
Poseidon is an efficient and scalable LLM training framework designed with heterogeneity awareness, which employs two efficient, theoretically grounded strategies: stage-level pruning via early stopping with partial estimation, and layer-to-stage mapping exploiting a ridge-like distribution pattern.
Xiao-Song Chen, Shao Nie, Zhong-Min Zhao et al.· 0 citations
Graph message passing offers a common way to express learning algorithms, physical simulations, and numerical solvers. Efficient execution depends on interaction structure and data movement, which can be obscured when a program is expressed as a sequence of tensor operations. On memory-constrained systems such as lapto...
Graph Neural Networks (GNNs) have become a fundamental tool for learning over graph-structured data. Under the message-passing framework, mainstream GNN models alternate between feature transformation and neighborhood aggregation. Fusing these two phases into a node-level pipelined push dataflow, in which each node’s t...
Shi Chen, Jun-Sheng Chang, Yang Guo et al.· ACM Transactions on Architec...· 0 citations
HDA-MoE is presented, a framework that optimizes MoE execution on NMP architectures through hybrid parallel deployment and runtime scheduling and integrates an offline hybrid parallel mapping algorithm with an online dynamic and adaptive scheduling mechanism to reduce communication overhead while improving computation...
Hao-Chen Huang, Shu-Zhang Zhong, Sheng-Xuan Qiu et al.· IEEE Transactions on Compute...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.