Shared Load-Store Unit for Instruction-Based Accelerators
Abstract
Modern systems-on-chip (SoCs) rely on heterogeneous accelerators for performance scaling. Memory access is a critical bottleneck, but the complexity of cache coherence and nuances of weak memory consistency models (which may vary from system to system) represent a significant designer burden to every load/store unit that must be created. This paper proposes a single shared load/store unit (LSU) to abstract away system details and offer a simplified interface for designers. We evaluate the shared LSU on an AMD ZCU104 FPGA platform using a soft RISC-V processor through two distinct case studies: (1) a homogeneous system using multiple vector processor cores, and (2) a heterogeneous system combining a vector core with a GEMM systolic array. Results demonstrate that the shared LSU reduces duplicated LSU logic and integration overhead while enabling significant performance gains, achieving a $\mathbf{2. 6 6} \times$ speedup in the homogeneous vector configuration over a single-vector baseline, and a $\mathbf{1 0. 9 3} \times$ speedup for neural network inference in the heterogeneous configuration over a scalar baseline.