Precision-Reconfigurable Zero-Skipping SRAM Compute-in-Memory Architecture for Energy-Efficient Edge-AI Acceleration
Abstract
The energy cost of data movement between static random-access memory (SRAM) and arithmetic units has become a critical limitation in edge artificial-intelligence accelerators. SRAM-based compute-in-memory (CIM) alleviates this bottleneck by executing multiply-and-accumulate operations near or inside the memory array. This work proposes a Precision-Reconfigurable Zero-Skipping SRAM-CIM (PRZS-CIM) architecture that combines bit-serial exact multiply-accumulate decomposition, 4/6/8-bit precision control, row-level zero gating, and all-zero bit-plane suppression. The method preserves exact arithmetic for the selected quantized representation because only mathematically zero terms are gated. A technology-independent architecture-level behavioral evaluation was performed using 20,000 randomly generated 64-element dot products per operating point, with activation sparsity varied from 0% to 80%. At 50% activation sparsity, the 8-bit mode preserves the full-precision numerical result while reducing the proposed cell-evaluation switching proxy by 79.55% relative to an ungated 8-bit bit-serial baseline. The 6-bit mode achieves 89.70% switching-proxy reduction with 3.36% normalized root-mean-square dot-product error relative to the original 8-bit operands, while the 4-bit mode reaches 95.62% reduction with 12.73% NRMSE. At 80% activation sparsity, switching-proxy reductions rise to 91.82%, 95.89%, and 98.25% for 8-, 6-, and 4-bit modes, respectively. The proposed architecture therefore offers a clear precision-energy trade-off while avoiding analog conversion overhead in the conceptual datapath. The reported proposed results are architecture-level simulation results rather than post-layout silicon measurements.