Skip to content

A Nonvolatile AI-Edge Processor With Lossless-Compressed-Computing STT-MRAM Near-Memory-Compute Macro Using Dynamic Floating-/Fixed-Point Accumulation

Oct 2026 · IEEE Journal of Solid-State Circuits · Vol 61, pp. 5937-5949 · 0 citations · 47 references

Abstract

Nonvolatile AI-edge processors based on near-memory-compute (nvNMC) enable energy-efficient multiply-and-accumulate (MAC) operations with short wakeup latency for edge inference operations. Lossless compression is required for floating-point (FP) neural network (NN) models under on-chip memory capacity constraints; however, this imposes several challenges: 1) data decompression overhead due to lossless FP weight encoding; 2) a read-yield–energy tradeoff due to near-far effects in large nonvolatile memory arrays; 3) a MAC-accuracy–energy tradeoff in FP accumulation within the multi-level adder tree; and 4) inefficient layer fusion due to feature-map (FM)–weight mapping in nvNMC. This article addresses these challenges by presenting a nonvolatile AI-edge processor featuring four system–circuit co-design schemes: 1) delta-exponent lossless compression computation (DELC ${}^{2}$ ); 2) near-far-aware bias-and-mapping (NFABM); 3) dynamic floating-/fixed-point (FXP) accumulation (DF2PA); and 4) nvNMC-friendly hybrid-layer-fusion weight-mapping (CIM-HLF-WM). A test chip fabricated using foundry-provided 22-nm spin-transfer-torque magnetic random access memory (STT-MRAM) achieved an energy efficiency of 24.6 TFLOPS/W and a wakeup-to-response latency of 428.58 $\mu $ s under BF16 input and weight, verifying its effectiveness for capacity- and energy-constrained AI-edge applications.

View source

Similar papers

Open access Sep 2026

Precision-Reconfigurable Zero-Skipping SRAM Compute-in-Memory Architecture for Energy-Efficient Edge-AI Acceleration

This work proposes a Precision-Reconfigurable Zero-Skipping SRAM-CIM (PRZS-CIM) architecture that combines bit-serial exact multiply-accumulate decomposition, 4/6/8-bit precision control, row-level zero gating, and all-zero bit-plane suppression.

Bitla Prabhakar T. Satyanarayana, Dr. Malothu Amru, Dr. Somala Rama Kishore · 0 citations
Review Open access Sep 2026

Nonlinear selectorless memory technologies for high-density storage and beyond-von-Neumann computing: a review

With the rapid advancement of artificial intelligence (AI), there is an urgent demand for emerging technologies that can deliver low-power storage alongside high-performance computing. Traditional volatile memories—such as static random access memory and dynamic random access memory—which are closely integrated with th...

Daphne Chen, M. Kozicki, Tuo-Hung Hou · 0 citations
#edge computing Preprint Sep 2026

FlexSpIM: An Event-Based Digital Compute-In-Memory Accelerator with Flexible Operand Resolution and Layer-Wise Hybrid Stationarity

FlexSpIM, a digital CIM architecture supporting arbitrary operand resolution and shape within a unified storage for weights and neuron states, is introduced, enabling a layer-level hybrid weight- and output-stationary dataflow, maximizing operand reuse and reducing costly on- and off-chip data movement during SNN execu...

Nicolas Chauvaux, Adrian Kneip, Charlotte Frenkel · 0 citations
Preprint Sep 2026

FALCON: Fault-Tolerant Magnetic Tunnel Junction-Based In-Memory Stochastic Architecture for Reliability-Critical Edge AI Applications

As modern data-centric applications such as neural inference and sensor-edge analytics expand, they increasingly encounter the von Neumann memory wall, suffering from excessive data movement overhead and stringent energy constraints. In-Memory Computing (IMC) utilizing emerging non-volatile technologies, such as Magnet...

Farzad Razi, M. Moghadam, Sercan Aygün et al. · 0 citations
Open access 2026

Sensing- and Periphery-Aware Modeling of a Hybrid MTJ–CMOS Nonvolatile 1-bit Sign-Weight Memory Subsystem for Energy-Constrained Edge Inference

The spin-transfer-torque magnetic random access memory (STT-MRAM) is attractive for nonvolatile 1-bit sign-weight storage, but a favorable cell or local-read metric does not establish subsystem efficiency. We analyze a hybrid magnetic tunnel junction (MTJ)–complementary metal–oxide–semiconductor (CMOS) memory subsystem...

Byeong-Gwon Kim, Byeong-Kwon Ju, Ki-Young Lee · 0 citations

Related blog posts

Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.