Shared Load-Store Unit for Instruction-Based Accelerators
Modern systems-on-chip (SoCs) rely on heterogeneous accelerators for performance scaling. Memory access is a critical bottleneck, but the complexity of cache coherence and nuances of weak memory consistency models (which may vary from system to system) represent a significant designer burden to every load/store unit th...