Neural Network-Oriented Chip Architecture Design for Edge AI: Co-Optimization of Efficient Algorithms and Low-Power Hardware
Abstract
Edge Artificial Intelligence (Edge AI) is increasingly used to deploy deep learning models on embedded devices. It reduces latency and improves privacy compared with cloud computing. However, edge devices are constrained by limited computation, memory, and power resources. As a result, efficient neural network deployment becomes a key challenge. This paper examines how to design neural network-oriented chip architectures under such constraints. In particular, it focuses on the co-optimization of efficient algorithms and low-power hardware. It analyzes the computational characteristics of neural networks and the constraints of edge devices. Meanwhile, it reviews typical accelerator architectures, including processing element arrays, dataflow design, and memory hierarchy. In addition, it discusses algorithm-level optimization methods such as quantization and pruning, as well as hardware-level low-power techniques. The results show that hardware-software co-design is necessary to achieve high performance and energy efficiency, and point to future research directions in Edge AI chip design.