Skip to content
Review Open access

AI Compilation and Hardware-Software Co-Design for Resource-Constrained Edge Devices: A Brief Review

Jul 2026 · Applied and Computational Engineering · Vol 247, pp. 137-144 · 0 citations

TL;DR

The reviewed work suggests that next-generation edge AI systems should focus on portable compiler toolchains, memory-aware optimization, energy-aware design, and standardized benchmarking methodologies for efficient inference execution on resource-constrained edge devices.

Abstract

Edge artificial intelligence (Edge AI) has emerged as a promising solution to provide real-time and privacy-aware intelligent services at proximity to data sources. However, it is difficult to implement deep learning on edge devices with limited resources, such as microcontrollers, embedded devices, and IoT terminals with low power consumption, because they have limited computing power, memory capacity, and energy supply. This paper introduces a concise overview on AI compilation and hardware co-design methodologies for efficient inference execution on resource-constrained edge devices. This paper provides an overview of the key challenges in deploying deep learning models on resource-constrained devices and presents illustrative solutions in terms of compression techniques and hardware implementation. The discussion includes graph optimization, operator fusion, quantization, memory-aware inference, hardware-specific code generation, accelerator-based inference, latency prediction, and runtime optimization. The reviewed work suggests that high-performance efficient edge AI requires cross-layer rather than isolated model-level compression. AI compilation enables more efficient execution by converting neural networks to optimized hardware-aware code, and hardware- co-design further improves latency, memory, and energy efficiency. These results imply that next-generation edge AI systems should focus on portable compiler toolchains, memory-aware optimization, energy-aware design, and standardized benchmarking methodologies.

Read PDF

Similar papers

Review Open access Jul 2026

Neural Network-Oriented Chip Architecture Design for Edge AI: Co-Optimization of Efficient Algorithms and Low-Power Hardware

Edge Artificial Intelligence (Edge AI) is increasingly used to deploy deep learning models on embedded devices. It reduces latency and improves privacy compared with cloud computing. However, edge devices are constrained by limited computation, memory, and power resources. As a result, efficient neural network deployme...

Yuchuan Yue · 0 citations
#federated learning Review Open access Sep 2026

Edge AI: Embedded Intelligence, a Review on Hardware and Applications

A unified analytical framework is introduced that models inference latency, bandwidth, energy, computational complexity, throughput, and model compression, and is used to compare cloud versus edge execution.

H. Al-Mimi, Ahmad Al-Dahoud, Ali A. Al-Dahoud et al. · 0 citations
Review Open access Aug 2026

Analysis of Research Progress on Deployment Methods for Deep Learning Models on FPGAs

A systematic review of FPGA-based DL deployment from a cross-layer perspective spanning model, compiler, architecture, runtime, and electronic design automation (EDA) is presented, highlighting that reliable cross-study comparison requires careful consideration of model configuration, precision, execution phase, batch...

Shuo Wang, Lei Chen, Chunsheng Tian et al. · 0 citations
#artificial intelligence Preprint Aug 2026

A Generalized Optimization Engine (GOE) for Edge AI Inference Acceleration

This study underscores the critical role of such a generalized optimization system in preparing model deployment over resource-constrained heterogeneous hardware in a tactical environment, and proposes a comprehensive hardware (HW) and model-agnostic generalized optimization architecture that integrates these technique...

V. Dasari, Jakob A. Adams, V. Mishra et al. · 0 citations
Open access Sep 2026

Versat-AI: An ONNX-to-SoC Compiler for Model-Agnostic CGRA Edge Inference

Edge inference on resource-constrained embedded nodes demands accelerators that are energy-efficient and compact. This paper presents Versat-AI, an open-source compiler that accepts a standard Open Neural Network Exchange (ONNX) model and generates a complete, synthesisable RISC-V System-on-Chip (SoC) with an embedded...

R. Teixeira, J. Rodrigues, Jaime Aguiar et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.