Skip to content

ThunderGNN: Unlocking Tensor Cores for Graph Neural Networks

Jun 2026 · Proceedings of the VLDB Endowment · Vol 19, pp. 2699-2712 · 0 citations · 49 references

TL;DR

ThunderGNN is a hardware-aware acceleration system designed to reconcile graph irregularity with Tensor Core rigidity and significantly outperforms state-of-the-art systems, achieving geometric mean speedups of 1.89× over DGL and 2.59× over PyG.

Abstract

Graph Neural Networks (GNNs) have emerged as the state-of-the-art methodology for learning on graph-structured data, yet their performance is severely constrained by a fundamental mismatch between irregular graph sparsity and the rigid parallelism of modern hardware. While modern GPUs rely on Tensor Cores (TCs) to deliver massive computational throughput, these units demand strictly tiled, dense inputs—a requirement that conflicts with the extreme sparsity of real-world graphs. Existing frameworks fail to resolve this design conflict: they either fallback to legacy SIMT cores, leaving TCs underutilized, or they incur prohibitive memory bloat by forcing sparse data into dense tiles via excessive padding. To bridge this gap, we propose ThunderGNN, a hardware-aware acceleration system designed to reconcile graph irregularity with Tensor Core rigidity. ThunderGNN employs a unified co-design strategy comprising three key optimizations: (1) a sparsity-aware reordering algorithm that logically groups graph rows to maximize local density; (2) a Condensed Binarized Abstraction (CBA) storage layout that physically organizes the adjacency matrix into TC-aligned blocks without explicit zero-padding; and (3) a hardware-aware execution engine that efficiently streams compressed blocks directly into TCs. Extensive experiments on NVIDIA A100 GPUs demonstrate that ThunderGNN significantly outperforms state-of-the-art systems, achieving geometric mean speedups of 1.89× over DGL and 2.59× over PyG.

View source

Similar papers

#graph neural networks Preprint Aug 2026

CoRe-GNN: Multilevel Message passing on Coarsened graphs

CoRe-GNN is proposed, which performs both propagations in parallel at each layer: a coarsened inter-cluster term capturing long-range structure, and a local intra-cluster term preserving per-node discriminability.

Antonin Joly, Nicolas Keriven, Aline Roumy · 0 citations
Book Open access Sep 2026

TileSpMM: A Variable-Size Tiled Algorithm for Sparse Matrix-Matrix Multiplication on Tensor Cores

This work proposes TileSpMM, which breaks the static-granularity bottleneck through a variable-size tiling algorithm that dynamically adapts to local sparsity patterns, and is equipped with an adaptive load-balancing strategy and customized granularity-specific kernels to improve hardware utilization and mitigate compu...

Hongwei Zeng, Shu-Qin Feng, Hao-Cheng Lian et al. · 0 citations
#graph neural networks Open access Sep 2026

PipeGNN: A Bandwidth-Efficient GNN Accelerator with Node-Level Pipelined Push Execution

Graph Neural Networks (GNNs) have become a fundamental tool for learning over graph-structured data. Under the message-passing framework, mainstream GNN models alternate between feature transformation and neighborhood aggregation. Fusing these two phases into a node-level pipelined push dataflow, in which each node’s t...

Shi Chen, Jun-Sheng Chang, Yang Guo et al. · 0 citations
Conference Open access May 2026

SNI-GNN: SmartNIC-Assisted Full-Graph GNN Training with In-Network Embedding Prediction

SmartNIC-based in-network prediction is a practical complement to partitioning and compression techniques for communicationefficient full-graph GNN training at scale, and results indicate SmartNIC-based in-network prediction is a practical complement to partitioning and compression techniques for communicationefficient...

Guofan Yu, Si-Tian Chen, Zhenheng Tang et al. · 0 citations
Preprint Aug 2026

Two-level domain-decomposition AdaGrad method for scalable training of graph neural networks

The proposed DD-AG2m alternates between AG2m optimization on the original (global) graph and AG2m optimization on the partitioned graphs, and introduces a two-level variant that performs global optimization steps on a coarse graph obtained by randomly subsampling nodes within each subdomain.

Laurynas Varnas, Julien Herrmann, Alexander Heinlein et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.