Jul 2026
Optimizing ML Workload Partitioning between CPUs and CIM Accelerators for Heterogeneous Computing
An Integer Linear Programming (ILP)-based workload partitioning framework for heterogeneous CPU-CIM systems that minimizes end-to-end inference latency under RRAM constraints, captures parallelism, and combines empirical profiling with analytical models is proposed.
Joel Klein, Rebecca Pelke, Roberto Laudani et al.
· arXiv.org · 0 citations