Skip to content
Open access

High-Throughput GPU Design and Implementation of MAYO with Matrix Computation Reordering

Sep 2026 · IACR Transactions on Cryptographic Hardware and Embedded Systems · Vol 2026, pp. 496-521 · 0 citations · 26 references

TL;DR

This work presents the first high-throughput GPU implementation and evaluation of MAYO, covering Keygen, Signing, and Verification, and introduces a data reuse-aware computation model for matrix-dominated and memory-access-bound workloads, and proposes Reuse Matrix Multiplication Evaluation (RMME) to reduce redundant global memory accesses through operand-level data reuse.

Abstract

With the rapid development of quantum computing, traditional digital signature schemes face increasing security threats. MAYO is a post-quantum digital signature scheme based on the hardness of the multivariate quadratic problem. Owing to its compact signature size, reduced public key, and flexible parameter sets, it has emerged as a competitive candidate in the NIST Additional Digital Signature round. However, existing MAYO implementations primarily target general-purpose processors, SIMD-enabled CPUs, or embedded platforms, and their throughput remains insufficient for high-demand scenarios such as server-side verification and largescale authentication services. In this work, we present the first high-throughput GPU implementation and evaluation of MAYO, covering Keygen, Signing, and Verification. We introduce a data reuse-aware computation model for matrix-dominated and memory-access-bound workloads, and propose Reuse Matrix Multiplication Evaluation (RMME) to reduce redundant global memory accesses through operand-level data reuse. Building upon this method, we further develop a GPU-oriented hardware– software co-design framework that combines algorithmic restructuring, warp-level kernel design, task-level memory reuse, and asynchronous multi-stream execution to maximize end-to-end throughput. Across all four MAYO parameter sets, the proposed implementation achieves more than 700x higher Signing throughput and more than 600x higher Verification throughput than CPU-based implementations. In the best case, it supports millions of verifications per second. These results demonstrate that RMME, together with GPU-oriented hardware–software co-design, is highly effective for accelerating multivariate signature schemes and provides a practical high-throughput solution for post-quantum digital signatures.

Read PDF

Similar papers

Sep 2026

Design Space Exploration of Accelerating Monolith Hash on Hardware for STARK

Zero-Knowledge Proof (ZKP) systems, particularly zk-STARKs, incur substantial computational overhead, where hash functions constitute a significant portion of the total cost. Among various ZK-friendly hash functions, Monolith achieves strong performance in both plaintext domain and ZK domain, making it a promising cand...

Cheng Chen, Gang-Qiang Yang, Hong-Chao Zhou et al. · 0 citations
Aug 2026

S2MM: Scalable FPGA Acceleration of Secure Matrix Multiplication with Homomorphic Encryption

Homomorphic Encryption (HE) enables secure computation on encrypted data, addressing privacy concerns in cloud computing. However, the high computational cost of HE operations, particularly matrix multiplication (MM), remains a major barrier to its practical deployment. Accelerating Homomorphic Encrypted MM (HE MM) is...

Zhi-Han Xu, Rajgopal Kannan, Viktor K. Prasanna · 0 citations
Preprint Sep 2026

FPGA Acceleration of Fully Homomorphic Encryption with Adaptive Key Switching

Fully Homomorphic Encryption (FHE) enables privacy-preserving cloud services but incurs substantial computation overhead, making hardware acceleration essential. Among FHE operations, key-switching is a major performance bottleneck. Recent cryptographic advances introduce a novel key-switching method (i.e., KLSS) that...

Zhi-Han Xu, Jayashree Adivarahan, Rajgopal Kannan et al. · 0 citations
Review Aug 2026

Hardware Acceleration of Fully Homomorphic Encryption: A Comprehensive Review of FPGA Implementations

Although Fully Homomorphic Encryption (FHE) enables computation over encrypted data, its substantial computational and storage overhead remains a major obstacle to practical deployment. Among available hardware platforms, FPGAs offer a favorable balance of performance, flexibility, and energy efficiency, making them a...

Lingyu Gong, Farhad Merchant · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.