Breaking the Area Limit of Composable Gadgets
Abstract
With the continuous advancement of side-channel attack techniques, the demand for high-order masking has been steadily increasing. However, as the masking order grows, the area overhead quickly becomes a critical bottleneck for cryptographic primitives in the IoT environment. Classical arbitrary-order masking schemes, such as HPC2 and HPC3, incur an intolerable O(d2) area overhead as the security order d increases, making high-order designs prohibitively costly in hardware. Even emerging approaches such as HO-TSM and CCHPC, which achieve O(d) latency, still retain area costs of the same magnitude. In contrast, recent optimization efforts, including COMAR and OBS, reduce the area cost but remain limited to low-order settings. In this work, we introduce Cycle-Mul (CM), a new composable hardware masking scheme built on a ring accumulation circuit. In detail, CM achieves the O(d) area overhead at the arbitrary security order d for the first time, demonstrating strong realization potential for high-order masking. We further optimize to reuse random sources in low-order gadgets, and reduce the second-order randomness cost to only two sources. At the circuit level, we propose a register-enable mechanism for the pipeline implementations and a fine-grained scheduling strategy for the clock-gating implementation. Leveraging the aforementioned optimizations, we evaluate the CM scheme on the AES S-box circuit masked with Trivium. Our schemes achieve area reduction of 17.1% (resp. 19.8%) at the second order and 57.3% (resp. 69.8%) at the fourth order under pipeline (resp. clock-gating) implementation compared to previous best-in-area schemes. For security, we validate our scheme using the formal verification tool PROLEAD and practical FPGA-based experiments, complementing the theoretical proof.