Skip to content

Author

Hella Toto Kiesa

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Sep 2026

Joint Exploration of Neural Networks and Systolic Hardware for Improved AI Accelerator Performance

When executing common Neural Networks (NNs) on custom AI accelerators, the high performance suggested by advertised Giga or Tera Operations per Second (GOPS/TOPS) is typically not achieved, as low hardware utilization often leads to an effective performance in the single-digit percentage range of the theoretical peak. This discrepancy arises as NNs are typically designed without accounting for the target hardware, leading to inefficient mappings and software optimizations that fail to deliver the expected gains. Addressing this, we present a hardware-aware workflow that combines accurate latency modeling and design space exploration to optimize both neural network architectures and the underlying systolic-array-based accelerator. We develop and validate two high-precision latency models for two different Row-Stationary (RS) dataflows on our target accelerator. Using these models together with a structured search space generation, we generate Pareto-optimal search spaces in terms of achieved GOPS and latency for a given hardware target, and use these for a Bayesian Bayesian Hardware-Aware Neural Architecture Search. We further explore the accelerator design itself in a subsequent hardware DSE stage, varying the PE array dimensions and clock ratio to identify hardware configurations that maximize efficiency and minimize inference latency for each network and dataflow. We demonstrate our approach on ResNet-like networks. On the original hardware, the discovered ResNet-50-like architecture achieves an 85% relative increase in hardware utilization, reduces latency by 21% and parameter count by 18%, and maintains baseline ImageNet accuracy. On the optimized hardware, latency and area are further reduced while efficiency is increased by up to 94%, with up to 20% fewer PEs compared to the baseline configuration. For ResNet-34, similar trends are observed, with latency reductions exceeding 33% and efficiency gains up to 43%. To enable reproducibility, we open-source our complete workflow of latency models, search space generation, NAS and training pipeline.

Annina Gutermann, Alexey Serdyuk, Foivos Paraskevas et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.