Aug 2026· International Symposium on Low Power Electronics and Design· pp. 1-7· 0 citations· 31 references
Computer Science
TL;DR
This work introduces a differentiable, GPU-accelerated CNN inference simulator that backpropagates through the analog accumulation trajectory, and allows gradient-based exploration of high-dimensional, heterogeneous hardware-parameter settings.
Abstract
Edge intelligence promises responsive, private, and energy-efficient sensing without continual dependence on remote compute. This demands convolutional neural network (CNN) accelerators that deliver substantially higher throughput and energy efficiency than conventional digital pipelines while preserving high accuracy. Time-domain analog CNN accelerators offer compact, high-throughput neural pipelines with reduced conversion and data-movement overhead, but also introduce modeling challenges. Hardware-aware models have been implemented in GPU-accelerated deep-learning libraries, but these are not designed for trajectory-dependent multiply-accumulate (MAC) operations where dynamic circuit behavior determines the final output. This work introduces a differentiable, GPU-accelerated CNN inference simulator that backpropagates through the analog accumulation trajectory. By making time-domain circuit dynamics compatible with automatic differentiation, the simulator allows gradient-based exploration of high-dimensional, heterogeneous hardware-parameter settings. Representative tuning studies show that layerwise tuning of time-dependent circuits can improve task-aware operating points rather than selecting a single global setting, and that gradient-based tuning outperforms a derivative-free baseline.
APEX is presented, a dual-sparsity SNN inference accelerator that integrates the PASC-IF neuron into the LoAS hardware framework, and guarantees mathematical equivalence between the converted SNN and the source ANN, thereby achieving ANN-equivalent accuracy at significantly reduced timesteps.
D. Venkatesh, S. Radhakrishnan, Rajshekhar Rakshit et al.· 0 citations
While Neural Radiance Fields (NeRF) have transformed 3D vision, their prohibitive computational and memory demands restrict real-time deployment on power-constrained edge devices. To bridge this gap, we propose a novel neural rendering accelerator that orchestrates three architectural innovations to maximize throughput...
Cheng Zhang, Yuefeng Zhang, Wenkai Zhou et al.· IEEE Transactions on Circuit...· 0 citations
Low-level hardware acceleration strategies to deconstruct the mapping from algorithm logic to silicon substrates are reviewed to provide strong guidelines for the hardware-software co-design of emerging ultra-low power edge Artificial Intelligence (AI) chips.
Linenxu Zhang· MATEC Web of Conferences· 0 citations
A systematic review of FPGA-based DL deployment from a cross-layer perspective spanning model, compiler, architecture, runtime, and electronic design automation (EDA) is presented, highlighting that reliable cross-study comparison requires careful consideration of model configuration, precision, execution phase, batch...
Shuo Wang, Lei Chen, Chunsheng Tian et al.· Electronics· 0 citations
PELSI incorporates a Genetic Algorithm to identify the near-optimal CPU-GPU layer-switched CNN inference configuration from within the large exponential design space that meets the given latency requirement most power efficiently.
PELSI incorporates a Genetic Algorithm to identify the near-optimal CPU-GPU layer-switched CNN inference configuration from within the large exponential design space that meets the given latency requirement most power efficiently.
Unknown authors· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.