Skip to content
Book Open access

StreamTrace: Fast Trace Analysis for Large-Scale Parallel Applications on a Single Node

Sep 2026 · Proceedings of the International Conference on Parallel Processing · pp. 671-681 · 0 citations · 13 references

Abstract

Trace-based performance analysis provides essential insights for understanding and optimizing large-scale parallel applications. However, traces from applications running on tens of thousands of processes can easily exceed terabytes, far surpassing the memory capacity of typical computing nodes. Existing approaches either require expensive distributed analysis or cannot work on platforms with limited memory. To address these challenges, we present StreamTrace, a streaming trace analysis system that enables efficient analysis of massive traces on a single computing node with bounded memory consumption. StreamTrace introduces two key techniques: communication pattern-aware chunk partitioning that minimizes cross-chunk dependence, and dynamic priority-based chunk scheduling that reduces analysis waiting time by prioritizing frequently depended upon processes. Our evaluation on traces from applications with up to 8,192 processes demonstrates that StreamTrace achieves up to 3.48 × speedup on a single node over existing distributed systems. The ablation study shows the proposed technique can improve the performance over naive streaming implementations by up to 8.50 ×.

Read PDF

Similar papers

Preprint Sep 2026

VarioPath: Workload-Aware All-to-All Communication for PCIe GPU Clusters

VarioPath, an efficient AlltoAllv scheduling framework for PCIe GPU systems that combines an offline topology-aware analyzer with an online demand-aware scheduler, and shows average AlltoAllv speedups of 5.88x over FAST and 1.72x over DeepEP.

Yao Fei, Jin Fang, Si-Ze Zheng et al. · 0 citations
Open access 2023

Performance Bottlenecks Necks in Data Heavy Python Applications

Performance improvements in data-intensive Python applications have become more critical due to the increasing computational needs of modern analytics, machine learning, and large-scale data processing systems. Although the Python environment is enormous as well as flexible in development, frequent delay in execution,...

Madhurima Kommuru, Appala Nooka Kumar Doodala · 0 citations
Book Open access Sep 2026

LMTracer: Fine-Grained and Real-Time Performance Profiling for Production LLM Systems

Training and serving large language models (LLMs) has become a core business for AI providers. To ensure a high-quality user experience while optimizing infrastructure costs, providers need to closely monitor the performance of LLM executions in production. However, existing performance profiling tools fall short in th...

Wei Liu, Yong-Chao He, Bo-Han Zhao et al. · 0 citations

Procyon: Promoting Fine-Grain Multi-Tenancy to Optimize Sparse Streaming Accelerators

Procyon, a fine-grain multi-tenancy framework that fuses the PE instruction streams of multiple workloads into a unified execution schedule, substantially reduces PE underutilization that results in 3 × speedup over state-of-the-art sparse streaming accelerators, and reaches a peak throughput of 61 .

Ubaid Bakhtiar, Jeremy Sha, Helya Hosseini et al. · 0 citations
Open access Aug 2026

Investigating Parallel Scaling Bottlenecks Across Rust, Julia, Haskell, and Python: Workload–Runtime Signatures

Parallel performance depends not only on programming language and runtime design, but also on how the dominant execution bottleneck changes as parallelism increases. We present a controlled cross-language study of Rust, Julia, Haskell, and Python using Merge Sort, Closest Pair of Points, and Numerical Sum in a multicor...

Muhammad Hassam Aslam Khan, Daniel Stapleton, Medha Kulkarni et al. · 0 citations
Preprint Oct 2026

Exploiting the Interplay of Compute- and Memory-Bound kernels in MPI Applications

Parallel applications are often designed for synchronous, lock-step execution, treating communication stalls as performance hazards. Yet, in a communication-light application without frequent synchronization points that alternates between compute-bound memory-bound execution, an MPI communication stall can act as an un...

Ayesha Afzal, Krishna Manda, G. Hager · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.