Sep 2026· IACR Transactions on Cryptographic Hardware and Embedded Systems· Vol 2026, pp. 496-521· 0 citations· 26 references
TL;DR
This work presents the first high-throughput GPU implementation and evaluation of MAYO, covering Keygen, Signing, and Verification, and introduces a data reuse-aware computation model for matrix-dominated and memory-access-bound workloads, and proposes Reuse Matrix Multiplication Evaluation (RMME) to reduce redundant global memory accesses through operand-level data reuse.
Abstract
With the rapid development of quantum computing, traditional digital signature schemes face increasing security threats. MAYO is a post-quantum digital signature scheme based on the hardness of the multivariate quadratic problem. Owing to its compact signature size, reduced public key, and flexible parameter sets, it has emerged as a competitive candidate in the NIST Additional Digital Signature round. However, existing MAYO implementations primarily target general-purpose processors, SIMD-enabled CPUs, or embedded platforms, and their throughput remains insufficient for high-demand scenarios such as server-side verification and largescale authentication services. In this work, we present the first high-throughput GPU implementation and evaluation of MAYO, covering Keygen, Signing, and Verification. We introduce a data reuse-aware computation model for matrix-dominated and memory-access-bound workloads, and propose Reuse Matrix Multiplication Evaluation (RMME) to reduce redundant global memory accesses through operand-level data reuse. Building upon this method, we further develop a GPU-oriented hardware– software co-design framework that combines algorithmic restructuring, warp-level kernel design, task-level memory reuse, and asynchronous multi-stream execution to maximize end-to-end throughput. Across all four MAYO parameter sets, the proposed implementation achieves more than 700x higher Signing throughput and more than 600x higher Verification throughput than CPU-based implementations. In the best case, it supports millions of verifications per second. These results demonstrate that RMME, together with GPU-oriented hardware–software co-design, is highly effective for accelerating multivariate signature schemes and provides a practical high-throughput solution for post-quantum digital signatures.
Zero-Knowledge Proof (ZKP) systems, particularly zk-STARKs, incur substantial computational overhead, where hash functions constitute a significant portion of the total cost. Among various ZK-friendly hash functions, Monolith achieves strong performance in both plaintext domain and ZK domain, making it a promising cand...
Cheng Chen, Gang-Qiang Yang, Hong-Chao Zhou et al.· ACM Transactions on Reconfig...· 0 citations
Homomorphic Encryption (HE) enables secure computation on encrypted data, addressing privacy concerns in cloud computing. However, the high computational cost of HE operations, particularly matrix multiplication (MM), remains a major barrier to its practical deployment. Accelerating Homomorphic Encrypted MM (HE MM) is...
Zhi-Han Xu, Rajgopal Kannan, Viktor K. Prasanna· ACM Transactions on Reconfig...· 0 citations
Although Fully Homomorphic Encryption (FHE) enables computation over encrypted data, its substantial computational and storage overhead remains a major obstacle to practical deployment. Among available hardware platforms, FPGAs offer a favorable balance of performance, flexibility, and energy efficiency, making them a...
Lingyu Gong, Farhad Merchant· ACM Transactions on Reconfig...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.