Skip to content
Preprint

MiCoPro: End-to-End Mixed Precision HW/SW Co-design with HW-aware Proxy Model

Aug 2026 · 0 citations · 42 references
Computer Science

TL;DR

The MiCo framework is proposed, a holistic MPQ exploration and deployment framework for edge AI applications that adopts a novel optimization algorithm to search for accuracy-optimal quantization configurations under strict latency constraints and is extended to MiCoPro, which introduces a robust Hardware-Aware Proxy model to enhance prediction accuracy and hardware versatility.

Abstract

Quantized Neural Networks~(QNN) with low-bitwidth data have proven promising in efficient storage and computation on edge devices. To mitigate accuracy degradation while maximizing speedup, layer-wise mixed-precision quantization~(MPQ) becomes a popular solution. However, existing algorithms for exploring MPQ schemes are limited in flexibility and efficiency. Comprehending the complex impacts of different MPQ schemes on post-training quantization and quantization-aware training results is a challenge for conventional methods. Furthermore, an end-to-end framework for the optimization and deployment of MPQ models is missing in existing work. To address these challenges, we propose the MiCo framework, a holistic MPQ exploration and deployment framework for edge AI applications. The framework adopts a novel optimization algorithm to search for accuracy-optimal quantization configurations under strict latency constraints. We further extended the framework to MiCoPro, which introduces a robust Hardware-Aware Proxy (HAP) model to enhance prediction accuracy and hardware versatility. By leveraging target-specific latency modeling, MiCoPro enables rapid exploration and direct deployment from PyTorch models to bare-metal C code. We demonstrate the versatility of our framework on both the BitFusion accelerator and SIMD-extended RISC-V processors, achieving up to 40\% of latency reduction with less than 3\% of accuracy drop.

View source

Similar papers

Aug 2026

Hardware-Aware Neural Network Deployment on Multi-Core in-Memory Computing Systems: A Compiler Perspective

Conventional compute-in-memory (CIM) deployment flows usually assume a fixed crossbar geometry, although DNN layers often have different channel counts and matrix shapes. The resulting shape mismatch leaves part of the array capacity unused and can increase the number of split-and-transfer operations during inference....

Kaiwen Deng, Sifan Sun, Hanjie Liu et al. · 0 citations
Preprint Aug 2026

Streamable Neural Video Compression: A Mixed Precision Approach for Cross-Platform Deployment

Neural Video Codecs (NVCs) offer unprecedented rate-distortion performance, making them highly attractive for bandwidth-constrained environments like 5G cellular networks and emerging satellite direct-to-cell (D2C) links. However, deploying NVCs in real-world streaming applications is severely hindered by cross-platfor...

Kasidis Arunruangsirilert, He-Ming Sun, J. Katto · 0 citations
Aug 2026

ADC-Free Compute-in-Memory for Error-Resilient and Energy-Efficient AI Accelerators

Analog compute-in-memory (CIM) has recently emerged as a novel paradigm for artificial intelligence compute, but the efficiency of CIM is heavily bottlenecked by the energy and area overhead of analog-to-digital (ADC). While replacing high-precision ADCs with 1-bit conversion significantly reduces peripheral overhead,...

Wei-Wei Zhao, Sohan Salahuddin Mugdho, Cheng Wang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Faster Than Flash: Exploiting Attention Sparsity for Efficient Long-Context Decoding

Faster Flash Decoding (FFD) is presented, a novel hardware-algorithm co-design framework designed to break the memory wall in long-context decoding and introduces the top-delta strategy, which dynamically filters blocks to achieve distribution-adaptive sparsity without global synchronization.

Zhigeng Liu, Zhiyuan Ning, Rui-Xiao Li et al. · 3 citations
#machine learning Preprint Sep 2026

OptiPrime: Optimizing Private Inference through Protocol-Hardware Co-design

Private deep neural network (DNN) inference based on hybrid homomorphic encryption (HE) and multi-party computation (MPC) can protect user data with a formal guarantee, but at the cost of significant latency overhead due to HE. Customized HE accelerators have been proposed and have achieved orders-of-magnitude speedup...

Jiang-Rui Yu, Ye Yu, Si Chen et al. · 0 citations
Open access 2026

PUMA: A PMU-Guided Multi-Domain Layer-Aware DVFS Framework for Low-Power Mobile AI on Smartphones

PUMA combines offline layer/layer-chain PMU characterization with online GPU PMU observations to identify execution characteristics associated with compute-bound, memory-bound, and bursty phases and achieves a lower energy-delay product than the existing governor across all evaluated workloads.

W. Chang, Seung-Ryeol Ohk, Young-Jin Kim · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.