Skip to content
Preprint

CIMERA: Compute-in-Interconnect and Memory with Reconfigurable Precision for LLM Inference

Jul 2026 · 0 citations · 12 references
Computer Science

Abstract

LLM impose significant computational and memory demands, creating challenges for energy-efficient inference across platforms ranging from data centers to power-constrained edge devices. Weight precision plays a critical role in balancing inference accuracy, throughput, and energy consumption, while modern LLM workloads exhibit pronounced heterogeneity and tolerance that favors adaptive precision execution. This paper presents CIMERA, a reconfigurable-precision LLM inference accelerator that integrates compute-in-interconnect and memory to mitigate the memory wall and enable precision-aware execution. Compared to Nvidia H100, CIMERA delivers up to $25\times$ and $10\times$ higher energy efficiency for 1B and 13B models, respectively.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.