Skip to content

Hardware-Aware Neural Network Deployment on Multi-Core in-Memory Computing Systems: A Compiler Perspective

Aug 2026 · IEEE Non-Volatile Memory System and Applications Symposium · pp. 1-6 · 0 citations · 17 references

Abstract

Conventional compute-in-memory (CIM) deployment flows usually assume a fixed crossbar geometry, although DNN layers often have different channel counts and matrix shapes. The resulting shape mismatch leaves part of the array capacity unused and can increase the number of split-and-transfer operations during inference. This paper presents FlexiCIM, a compiler framework for neural network deployment on multi-core CIM systems. FlexiCIM groups physical cores into virtual computing units (VCUs), maps reshaped layer weights under CIM capacity constraints, and schedules dependent tasks while accounting for communication overhead. The framework also uses an evolutionary search procedure to select VCU partitions for latency-first or throughput-first objectives. Experiments on five CNN models show that FlexiCIM achieves an average utilization of 75.5% under the default setting. In the latency-first setting, FlexiCIM reduces latency by up to 54.8% and obtains the lowest normalized energy among the evaluated designs (1.00 vs. 2.02/1.37/1.60 for Fixed-L/M/S). In the throughput-first setting, FlexiCIM provides the highest throughput on all five models, with $1.07\times-1.38\times$ speedup over the best fixed-size baseline. These results indicate that compiler-managed VCU partitioning and capacity-aware mapping improve deployment efficiency on multi-core CIM architectures.

View source

Similar papers

2026

Automatic Model Compression and Quantized Deployment of Convolutional Neural Networks on Programmable Data Planes

The rapid development of programmable network devices and the widespread adoption of machine learning (ML) in networking have facilitated efficient research into intelligent data planes (IDPs). Offloading ML to programmable data planes (PDPs) enables quick analysis and responses to network traffic dynamics, and efficie...

Xiaoquan Zhang, Mai Zhang, L. Cui et al. · 0 citations
Oct 2026

SynergyScale: Optimizing Offloading and Task Partitioning for Efficient Model Training

Deep neural networks (DNNs) with billions of parameters power many important applications, but their training is fundamentally constrained by the limited on-chip memory of GPUs. This memory wall forces training to rely on distributed execution or memory offloading, both of which introduce substantial inefficiencies. Ex...

Xiaoyang Sun, Jie Xu, Zheng Wang · 0 citations
Aug 2026

A Flexible Framework for Layer-Parallel CNN Training on FPGA Clusters

We present a flexible and scalable hardware/software framework for training CNNs on Ethernet-connected FPGA clusters using tightly pipelined layer parallelism. Starting from a high-level CNN and cluster description, the system automatically maps layers onto (potentially heterogeneous) devices, generates FPGA-specific b...

Philipp Kreowsky, Justin Knapheide, B. Stabernack · 1 citation
Aug 2026

Learning to schedule: interference-aware optimization for edge AI inference on shared GPUs

A time-varying integer program to minimize the long-term total cost of the edge AI inference system, including the inference latency, the inference error rate, the query-dispatching communication cost, and the energy consumption, subject to resource and workload constraints is proposed.

Ming-Tao Ji, Hehan Zhao, Lei Jiao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.