Area–Power–Performance Trade-Offs in Lightweight Non-Pipelined and Pipelined NTT Accelerators for CRYSTALS-Dilithium
This article investigates two lightweight field-programmable gate array (FPGA) implementations of an iterative NTT-based polynomial multiplication accelerator through non-pipelined and 4-stage pipelined architectures, showing that the non-pipelined architecture provides reduced hardware overhead and lower power consumption, whereas the pipelined architecture improves timing scalability and successfully operates at 280 MHz.