Oct 2026· IEEE Journal of Solid-State Circuits· Vol 61, pp. 5937-5949· 0 citations· 47 references
Abstract
Nonvolatile AI-edge processors based on near-memory-compute (nvNMC) enable energy-efficient multiply-and-accumulate (MAC) operations with short wakeup latency for edge inference operations. Lossless compression is required for floating-point (FP) neural network (NN) models under on-chip memory capacity constraints; however, this imposes several challenges: 1) data decompression overhead due to lossless FP weight encoding; 2) a read-yield–energy tradeoff due to near-far effects in large nonvolatile memory arrays; 3) a MAC-accuracy–energy tradeoff in FP accumulation within the multi-level adder tree; and 4) inefficient layer fusion due to feature-map (FM)–weight mapping in nvNMC. This article addresses these challenges by presenting a nonvolatile AI-edge processor featuring four system–circuit co-design schemes: 1) delta-exponent lossless compression computation (DELC ${}^{2}$ ); 2) near-far-aware bias-and-mapping (NFABM); 3) dynamic floating-/fixed-point (FXP) accumulation (DF2PA); and 4) nvNMC-friendly hybrid-layer-fusion weight-mapping (CIM-HLF-WM). A test chip fabricated using foundry-provided 22-nm spin-transfer-torque magnetic random access memory (STT-MRAM) achieved an energy efficiency of 24.6 TFLOPS/W and a wakeup-to-response latency of 428.58 $\mu $ s under BF16 input and weight, verifying its effectiveness for capacity- and energy-constrained AI-edge applications.
This work proposes a Precision-Reconfigurable Zero-Skipping SRAM-CIM (PRZS-CIM) architecture that combines bit-serial exact multiply-accumulate decomposition, 4/6/8-bit precision control, row-level zero gating, and all-zero bit-plane suppression.
Bitla Prabhakar T. Satyanarayana, Dr. Malothu Amru, Dr. Somala Rama Kishore· International Journal of Adv...· 0 citations
With the rapid advancement of artificial intelligence (AI), there is an urgent demand for emerging technologies that can deliver low-power storage alongside high-performance computing. Traditional volatile memories—such as static random access memory and dynamic random access memory—which are closely integrated with th...
Daphne Chen, M. Kozicki, Tuo-Hung Hou· Nanotechnology· 0 citations
FlexSpIM, a digital CIM architecture supporting arbitrary operand resolution and shape within a unified storage for weights and neuron states, is introduced, enabling a layer-level hybrid weight- and output-stationary dataflow, maximizing operand reuse and reducing costly on- and off-chip data movement during SNN execu...
Nicolas Chauvaux, Adrian Kneip, Charlotte Frenkel· 0 citations
As modern data-centric applications such as neural inference and sensor-edge analytics expand, they increasingly encounter the von Neumann memory wall, suffering from excessive data movement overhead and stringent energy constraints. In-Memory Computing (IMC) utilizing emerging non-volatile technologies, such as Magnet...
Farzad Razi, M. Moghadam, Sercan Aygün et al.· 0 citations
The spin-transfer-torque magnetic random access memory (STT-MRAM) is attractive for nonvolatile 1-bit sign-weight storage, but a favorable cell or local-read metric does not establish subsystem efficiency. We analyze a hybrid magnetic tunnel junction (MTJ)–complementary metal–oxide–semiconductor (CMOS) memory subsystem...
Byeong-Gwon Kim, Byeong-Kwon Ju, Ki-Young Lee· IEEE Journal on Exploratory...· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 8, 2026
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.