Skip to content
Preprint

Memory AMP: Overflow Avoidance, Complexity Reduction, and Comparative Analysis

Aug 2026 · 0 citations · 46 references
Computer Science Engineering Mathematics

TL;DR

A general gradient-based formulation for designing MAMP algorithms is developed and it is shown that the computation of the orthogonalization parameters in this formulation can suffer from catastrophic cancellation, which explains the finite-precision instability of WS-CG-VAMP.

Abstract

Approximate message passing (AMP)-type algorithms are widely used for signal recovery in high-dimensional noisy linear systems. Recently, a framework called memory AMP (MAMP) was introduced, offering a new approach to incorporating memory terms within AMP algorithms. Building on this, a low-complexity gradient descent MAMP (GD-MAMP) was proposed for right-unitarily invariant matrices. In this paper, we first address an overflow problem in GD-MAMP caused by intermediate variables exceeding the floating-point range, which typically occurs when the condition number is large. Second, we propose two low-complexity variants of GD-MAMP: one replaces full-length memory with partial memory, while the other reduces the number of matrix-vector products per iteration by $1/3$ (from three to two). Neither degrades the convergence speed notably. Third, we develop a general gradient-based formulation for designing MAMP algorithms. This formulation recovers warm-started conjugate gradient VAMP (WS-CG-VAMP) as a special case. Furthermore, we show that the computation of the orthogonalization parameters in this formulation can suffer from catastrophic cancellation, which explains the finite-precision instability of WS-CG-VAMP. Finally, we derive an equivalent reformulation, termed WS-CG-VAMP(r), which reduces the number of matrix-vector products by up to $50\%$. Measured by matrix-vector products, GD-MAMP converges faster for small condition numbers, whereas WS-CG-VAMP(r) converges faster for large ones under high-precision arithmetic but may diverge in IEEE double precision due to catastrophic cancellation.

View source

Similar papers

Aug 2026

ADC-Free Compute-in-Memory for Error-Resilient and Energy-Efficient AI Accelerators

Analog compute-in-memory (CIM) has recently emerged as a novel paradigm for artificial intelligence compute, but the efficiency of CIM is heavily bottlenecked by the energy and area overhead of analog-to-digital (ADC). While replacing high-precision ADCs with 1-bit conversion significantly reduces peripheral overhead,...

Wei-Wei Zhao, Sohan Salahuddin Mugdho, Cheng Wang et al. · 0 citations
Book Open access Aug 2026

MITRA: Reconfigurable, Low-Latency, and Power-Efficient In-Memory Stochastic Architecture for Transcendental Functions

Processing in memory (PIM) offers a compelling pathway to overcome the data movement bottleneck in modern AI and data-centric systems. This work introduces MITRA, a reconfigurable magnetic tunnel junction (MTJ)-based in-memory architecture that leverages stochastic computing (SC) to implement a broad class of transcend...

Farzad Razi, M. Moghadam, M. Najafi et al. · 0 citations
Preprint Aug 2026

Arithmetic Variable LogLog: Advancing the Memory-Variance Frontier

Cardinality estimation - counting the number of distinct elements in a data stream - requires a tradeoff between memory and accuracy. ExaLogLog recently established the state of the art for this tradeoff by combining wide registers with a Fisher-information-optimal maximum likelihood (ML) estimator, achieving the best...

Brian Bushnell · 0 citations
#small language model Preprint Aug 2026

Ternary-Valued Finite-Difference Time-Domain Method: Equivalence with the Yee Scheme Through Noise-Shaped Quantisation

It is demonstrated that finite-difference time-domain dynamics can be reproduced with field variables restricted to the ternary alphabet, and its extension to acoustics, Virieux-type elastodynamics and Schrodinger-equation solvers points to a broader class of quantised physics solvers for resource-constrained and speci...

I. S. Maksymov · 0 citations
Open access Aug 2026

RRAM Circuit-Enabled Nonlinear Precoding and Bit Precision Analysis

The rising number of users and antennas imposes exponentially growing computational loads on future communication systems. Yet conventional processors are facing a bottleneck for their nature of memory-computing separation. In-memory computing (IMC) emerges as a promising solution leveraging its intrinsic high parallel...

Yu-Hao Zhang, Hai-Fan Yin, Tao Wang et al. · 0 citations
Preprint Sep 2026

High-frequency Multispeculative Multiply-Accumulation Unit for Fused Posit Arithmetic

Posit arithmetic offers a compelling alternative to the IEEE 754 floating-point standard, providing enhanced accuracy. Its fused multiply-accumulate operations avoid intermediate rounding, ensuring exact numerical reproducibility through the quire, a wide fixed-point accumulator spanning the format's full dynamic range...

Mario Alonso, Miguel A. Sacristán, G. Botella et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.