Skip to content
Open access

A High-Speed, Low-Power Pipelined CNN Implementation on FPGA Using On-Chip Processing and Layer-Wise Streaming Control

2026 · IEEE Access · Vol 14, pp. 113658-113681 · 0 citations · 30 references
Computer Science

Abstract

Convolutional Neural Networks (CNNs) are extensively used in advanced image processing applications. However, their computational complexity makes real-time deployment challenging. In this research, a high-speed CNN classification model with full layer-wise control is proposed, enabling seamless streaming of pixel data across all CNN layers via line buffers. The proposed architecture employs a novel, optimized, pipelined, layer-wise streaming control mechanism that enables concurrent execution across multiple layers and achieves ultra-high processing speed suitable for real-time applications. The learnable parameters (weights and biases) obtained after training the Optimized CNN (O-CNN) model are used to design the O-CNN model’s hardware architecture. These parameters are stored on-chip to eliminate frequent fetching of data from memory and thereby improve speed while reducing power consumption. Furthermore, line-buffer-based dataflow and on-chip storage of fixed hardware parameters minimize memory access overhead, resulting in high classification speed, low-power, and low energy consumption per classification. The operations of different layers, clock cycles requirements, and timing behaviour are verified through HDL simulation. The proposed implementation uses a signed 32-bit, Q18.13 fixed-point number format and is implemented on the Xilinx ZCU104 FPGA board. With 796 cycles at a 15 ns clock (66.67 MHz) and a positive worst slack of 1.439 ns, the design achieves a processing speed of $11.94~\mu $ s per image while consuming only 1.745 W of power.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.