Aug 2026· Journal of Circuits, Systems and Computers· 0 citations· 25 references
Computer Science
TL;DR
This paper presents a novel memory controller architecture and a RISC-V instruction set extension to optimize MLC NVM write operations by balancing speed and retention time, and introduces a fast-store instruction in RISC-V to increasing write performance while addressing retention limitations.
Abstract
Non-volatile memory (NVM) technologies, particularly Multi-Level Cell (MLC) NVMs, offer significant potential for increasing memory density. MLC NVMs provide a tradeoff between write latency and retention time, where faster writes/stores result in lower retention and slower writes yield higher retention. However, limited work has been done to validate and prototype NVM-based systems in hardware, leveraging this tradeoff at the system level.
In this paper, we present a novel memory controller architecture and a RISC-V instruction set extension to optimize MLC NVM write operations by balancing speed and retention time. Our custom NVM controller, built around a finite state machine with an AXI memory-mapped interface, efficiently manages read/write operations with enhanced burst transfers, minimizing latency. Additionally, we introduce a fast-store instruction in RISC-V to increasing write performance while addressing retention limitations. Further, we design a dedicated AXI slave peripheral that supports bit-significance-aware writes: critical bits (e.g., MSBs) are written using slower, high-retention writes, while non-critical bits (e.g., LSBs) use faster, low-retention writes to help enhance performance without compromising data reliability. These enhancements are implemented in hardware on an FPGA platform. Experimental results show that our controller reduces hardware overhead by 30% compared to conventional designs, and the fast-store instruction improves performance by over 7% for streaming workloads with less than 0.08% hardware overhead. The bit-wise AXI peripheral has a LUT utilization staying below 3.5% even for 64×64 matrices, and under 1% for 32×32 sizes, making it viable for integration into larger SoCs.
Compared to conventional volatile memory, non-volatile memory (NVM) technologies offer lower static power, higher density, and data retention. However, they induce high write overheads in terms of energy and potentially latency. Hybrid cache architectures mix volatile and non-volatile technologies, with the idea to kee...
Stefan Meißner, Stefan Wildermann, Neele Peter et al.· Proceedings of the 4th Works...· 0 citations
SerMC is proposed, a memory-controller-based tripwire mechanism that validates accesses when metadata arrives from Dynamic Random Access Memory and extends enforcement to DRAM-bound accesses issued by both processors and DMA-capable devices, and provides practical average-case overhead but still exposes a clear worst-c...
Gen Xu, Li Lv, Jiayan Dong· Computers, Materials & C...· 0 citations
Modern systems-on-chip (SoCs) rely on heterogeneous accelerators for performance scaling. Memory access is a critical bottleneck, but the complexity of cache coherence and nuances of weak memory consistency models (which may vary from system to system) represent a significant designer burden to every load/store unit th...
Joseph Maheshe, G. Lemieux· IEEE International Conferenc...· 3 citations
The key-value (KV) cache is the dominant memory bottleneck in long-context large language model (LLM) decoding: every step reads it entirely, so decoding is memory-bandwidth bound. Holding a quantized KV cache in dense on-chip non-volatile memory (NVM) removes the off-chip transfer. Existing KV quantization methods, ho...
Jia-Hao Zheng, Yi-Fan Qin, Xiaobo Sharon Hu et al.· 0 citations
In-storage computing (ISC) is considered a next-generation memory architecture for its ability to relieve the data bottleneck between the host and the memory. While the required resources of large language models (LLMs) have increased significantly in recent years, the memory density has not scaled accordingly. Recentl...
Sanghun Shin, Sangyeon Kim, Gisan Ji et al.· 0 citations
Evaluation using cycle-accurate hardware-software co-simulation demonstrates that CHIPSMORE effectively sustains high throughput across varying model sizes, context lengths, and batch sizes while maintaining favorable power scaling.