Skip to content
Preprint

DiffPower: GPU-Accelerated Differentiable Switching Power Analysis and Optimization

Aug 2026 · 0 citations · 29 references
Computer Science

TL;DR

DiffPower translates design netlists into a PDK-agnostic bytecode representation, enabling analytical gradient computation via reverse-mode automatic differentiation, achieving up to a speedup over single-threaded CPU propagation on the largest evaluated design, with the GPU advantage growing with design scale.

Abstract

Accurate and scalable switching power analysis remains a critical bottleneck in modern physical design, often forcing a trade-off between computational speed and modeling fidelity. We present DiffPower, a GPU-accelerated framework for differentiable power analysis and optimization. DiffPower translates design netlists into a PDK-agnostic bytecode representation, enabling analytical gradient computation via reverse-mode automatic differentiation, achieving up to a $1{,}002\times$ speedup over single-threaded CPU propagation on the largest evaluated design, with the GPU advantage growing with design scale. A hybrid propagation methodology fusing analytical modeling with parallel simulation achieves a median toggle-rate correlation of $r{=}0.96$ across ten industrial and benchmark designs. The resulting \emph{power gradients}, computed up to $904\times$ faster than CPU finite-difference methods with near-perfect rank agreement, enable two downstream applications: (1) gradient-weighted cell sizing, which achieves up to $2.98\times$ improvement over local-power heuristics on industrial designs, with even stronger advantages at the 117K-cell scale where competing methods plateau; and (2) power virus generation via gradient ascent, which yields up to $2.13\times$ higher transition-weighted power, replacing a search process that traditionally requires hours.

View source

Similar papers

Review Open access 2026

A Review of GPU-Accelerated Finite-Difference Numerical Simulation: Four-Phase Synthesis, Design Rules, and a Metrics Card

This review synthesizes research on graphics processing unit (GPU)-accelerated finite-difference numerical simulation (FDNS) from 2003 to 2025 to clarify how GPU computing has reshaped simulation workflows rather than merely accelerating isolated numerical kernels and to address challenges related to comparability, ver...

Jiaxiang Liu, Chengpu Peng, Xian-Zhang Ling · 0 citations
Conference Open access Sep 2026

Analyzing the Impact of Architectural Design Decisions on Performance Across Generations of NVIDIA GPUs

The high-performance computing industry is moving beyond an era in which each generation of GPU provides uniform performance gains across all applications. The growing importance of AI is driving GPU architecture towards greater specialization, with more silicon devoted to Tensor Cores and reduced-precision arithmetic....

Matthew Tindale, I. Karlin, Tobias Salamon et al. · 0 citations
Jul 2026

FSZ: Breaking the Prediction-Throughput Trade-off in GPU Lossy Compression

FSZ, a GPU error-bounded lossy compressor that redesigns the prediction stage with three mutually reinforcing algorithmic innovations to achieve both higher compression ratios and higher throughput within a single CUDA kernel, achieves the highest average throughput among all evaluated compressors.

Jiajun Huang · 0 citations
Jul 2026

PortLBM: A Portable Lattice Boltzmann Tool Leveraging SYCL on AMD, NVIDIA, and Intel GPUs

The lattice Boltzmann method (LBM) is a well-established approach for simulating fluid flows at the mesoscopic scale. With the stagnation of Moore's law, high-performance computing has shifted toward GPU accelerators, necessitating programming models that ensure both portability and efficiency across diverse hardware p...

Alexander Strack, Marcel Graf, Alexander Van Craen et al. · 0 citations
Preprint Sep 2026

GPU-Accelerated Gate-Level Time-Based Power Analysis via Event-Density-Aware Partitioning and Kernel Fusion

Power analysis is crucial in modern chip design flow. Particularly, time-based power analysis can provide fine-grained power consumption information to facilitate the diagnosis of power issues and guide power optimization accordingly. However, it may take tens of hours to conduct time-based power analysis on modern lar...

Wei-Hao Wang, Yi-Kang Ouyang, Hong-Yuan Liu et al. · 0 citations
Preprint Aug 2026

FaCTz: Fast Critical-Point and Topology-Aware GPU Compression for Scientific Vector Fields

Error-bounded lossy compression is essential for storing and transferring the vector-field data produced by large-scale scientific simulations. Although it enforces a user-specified error bound to limit numerical distortion, it does not preserve the field's topology: small admissible perturbations can create or elimina...

Mingze Xia, Yu-Xiao Li, Sheng Di et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.