Skip to content
Preprint

PEAT: Pseudo-Error Assessment for GPU Kernel Validation in DNN Training

Sep 2026 · 0 citations · 63 references
Computer Science

TL;DR

The applicability of PEAT is demonstrated by presenting the results and analysis using GPUs from the two most popular vendors, NVIDIA V100 and AMD MI250, on various AI models, from vision tasks to language models, for both pretraining and finetuning scenarios.

Abstract

Deep neural networks (DNNs) are widely adopted in various fields, driving an emerging trend in developing software stacks associated with DNN training systems. For example, many codes have been ported across different frameworks or developed to leverage the computing power of GPUs or domain-specific accelerators. However, validating a kernel implementation in DNN training is time-consuming and generally requires massive storage. Specifically, this poses a fundamental question: how to characterize the behavior of a new implementation when it is integrated into a DNN training flow. Unfortunately, this problem is not well investigated in the literature, to the best of our knowledge. To address this shortcoming, we present PEAT - a lightweight inspection framework for \underline{P}seudo-\underline{E}rror \underline{A}ssessment associated with GPU kernel validation in DNN \underline{T}raining. Firstly, inspired by conventional fault injection (FI), PEAT's Profiler invokes an operation-wise kernel in a training flow to collect a DNN model's states (e.g., checkpoints and activations). More importantly, the Profiler introduces two simple yet effective techniques, playback FI and frequency-based runtime FI, leveraging persistent kernel calling during the training process. Secondly, PEAT's Analyzer characterizes profiled errors, revealing some signatures from the error distribution of a kernel compared to the golden one. Lastly, PEAT's Detector provides some guidelines as a sufficient condition, which enables associating several well-known error models with signature patterns. We demonstrate the applicability of our approach by presenting the results and analysis using GPUs from the two most popular vendors, NVIDIA V100 and AMD MI250, on various AI models, from vision tasks to language models, for both pretraining and finetuning scenarios.

View source

Similar papers

#machine learning Preprint Sep 2026

Accelerating the Mitigation of LLM Inference Nondeterminism Across GPU Architectures

Large language model (LLM) outputs are expected to be reproducible under greedy decoding, yet in practice the same model, prompt, and software stack produce different outputs on different GPUs. The root cause is floating-point non-associativity combined with hardware-dependent kernel selection. Inference frameworks sel...

L. Cooper, Shinnung Jeong, Hyeran Jeon et al. · 0 citations
#machine learning Preprint Sep 2026

KernelZero: Co-Evolving Proposer and Coder for Continuously Improved GPU Kernel Generation

High-performance GPU kernels are essential to modern machine learning systems, yet automatically generating kernels that are both correct and efficient remains challenging. Existing LLM-based approaches face two major limitations: the scarcity of high-quality training data aligned with the model's current capabilities,...

Changxin Ke, Rui Zhang, Zi-Xiang Fang et al. · 0 citations
Preprint Sep 2026

FINN-Tro: Exploiting Verification Gaps in Dataflow Inference Accelerators

The growing adoption of dataflow accelerators for neural network inference introduces new attack surfaces that existing verification methodologies fail to address. Inference ac- celeration frameworks such as FINN, which transform quantized neural networks into FPGA-deployable dataflow architectures, implicitly assume s...

Qazi Arbab Ahmed, Suraj Karki, Thorsten Jungeblut · 0 citations
#natural language process... Preprint Sep 2026

AMDKernelVault: Large-Scale Datasets and Agentic Training for AMD GPU Kernel Optimization

We introduce AMDKernelVault, an open HIP and Triton kernel corpus and training framework for recent AMD CDNA GPUs. Existing LLM-based kernel agents are largely CUDA/NVIDIA-centric and often depend on repeated frontier-LLM calls for generation, reflection, and optimization. To address this gap, we develop HIPKernelGen a...

Ji Liu, S. Majumder, Yi-Qing Huang et al. · 0 citations
#machine learning Preprint Sep 2026

AttnFuse: A Composable DSL for Compiling Attentions to Fused GPU Kernels

Modern AI systems are built on the Transformer architecture, whose core operation, attention, accounts for the majority of computation and memory cost. Researchers continually propose new attention variants to improve quality, efficiency, or context length, but each variant currently requires expert-written GPU code to...

Varun Kumar Dasoju, Tian Zhao · 0 citations
Conference Aug 2026

Designing and Building an FPGA Accelerator That Uses Less Energy for DNN Inference

Deep Neural Networks (DNNs) are critical to modern AI applications, yet their deployment on standard CPUs and GPUs is constrained by high power consumption and computational latency, particularly in resource-constrained edge environments. To address these limitations, this paper presents the design and implementation o...

P. V. G. K. Rao, Dudekula Raziya · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.