Modern processors increasingly provide matrix engines whose peak arithmetic throughput greatly exceeds conventional SIMD, but scientific applications rarely realize this advantage end to end. We examine this gap in SPECFEM3D's dominant stiffness operator on the Arm LX2 CPUs that power the flagship Lineshine supercomput...
Yinuo Wang, Lin Gan, Tianqi Mao et al.· 0 citations
StreamTrace is presented, a streaming trace analysis system that enables efficient analysis of massive traces on a single computing node with bounded memory consumption and introduces two key techniques: communication pattern-aware chunk partitioning that minimizes cross-chunk dependence and dynamic priority-based chun...
Yu-Yang Jin, Ji-Dong Zhai· Proceedings of the Internati...· 0 citations
Trace-based performance analysis provides essential insights for understanding and optimizing large-scale parallel applications. However, traces from applications running on tens of thousands of processes can easily exceed terabytes, far surpassing the memory capacity of typical computing nodes. Existing approaches eit...
Yuyang Jin, Ji-Dong Zhai· Proceedings of the Internati...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.