A 22-nm End-to-End Edge--AI Processor With Booth-Value-Confined Acceleration and Hardware-Aware Layer-Wise Model Deployment
Edge devices capable of running artificial intelligence (AI) applications have seen a surge in demand for energy-efficient and high-throughput computation. In this study, a 22-nm edge–AI processor, incorporating an accelerator with error-free Booth-value-confined (BVC) multiprecision (MP) multiplier and near-memory computing (NMC), is introduced to accelerate neural networks (NNs). It has the following three major features. First, a BVC MP multiplier based on radix-8 Booth (R8B) is introduced to reduce computation complexity by prohibiting the “±3” cases and support error-free training on GPU without accuracy loss originating from the mismatch between training and deployment. A PE is built based on this multiplier for parallel computation with 82% power reduction and 70% area reduction. Second, the proposed NMC-friendly data flow supports efficient data reuse and hence reduces off-chip memory traffic. The data flow supports data reuse of up to 16 times, matching the number of PEs and enabling regular read and write patterns. Third, a hardware-aware layer-wise model deployment approach is proposed with a memory space contiguity-aware (MSCA) model reshape strategy, and a hardware-aware NN splitting and scheduling algorithm. The proposed MSCA strategy maximizes burst access, and the proposed algorithm achieves efficient computation with high data reuse and low memory access. This deployment approach can achieve a reduction in memory access latency of 16.6%–32.0%. Measurements on a 22-nm test chip demonstrate a peak power efficiency of 33.98 TOPS/W under synthetic full-PE-utilization conditions, while achieving 12.92–29.11 TOPS/W for end-to-end NN inference on DarkNet19, VGG16, ViT-Tiny, and ResNet34.