Aug 2026· ACM Transactions on Design Automation of Electronic Systems· 0 citations· 10 references
TL;DR
SynaSpace is a behavior-driven configuration optimization framework for fault detection in logic synthesis tools that focuses on synthesis behavior coverage to guide configuration search, by constructing behavioral representations through joint analysis of synthesis logs and gate-level netlists.
Abstract
As FPGA design complexity increases, the correctness and reliability of logic synthesis tools are critical to ensuring correct hardware implementation. These tools translate hardware description languages (e.g., Verilog) into gate-level netlists, where latent faults may introduce functional errors or performance degradation during synthesis. Existing approaches rely on automatically generated Verilog test cases to find these latent faults. However, their effectiveness depends heavily on generator configurations and is typically guided by input diversity, which fails to accurately capture differences in synthesis behavior. Moreover, the high-dimensional configuration space of generator further hinders efficient exploration. To address these challenges, we propose SynaSpace, a behavior-driven configuration optimization framework for fault detection in logic synthesis tools. SynaSpace focuses on synthesis behavior coverage to guide configuration search, by constructing behavioral representations through joint analysis of synthesis logs and gate-level netlists. The framework comprises four components: (1) configuration space modeling for unified parameter representation; (2) Bayesian optimization–based configuration search for efficient exploration; (3) synthesis behavior characterization and coverage evaluation for capturing and quantifying behavioral differences; and (4) fault detection and utility modeling for extracting effective feedback via differential testing and deduplication. These components are integrated into a unified optimization framework to enable efficient configuration exploration and improved testing effectiveness. We evaluate SynaSpace on two established logic synthesis tools (i.e., Vivado and Yosys). SynaSpace identifies 18 faults across four categories, all of which have been confirmed and fixed by vendors and the open-source community.
Can large language models generate not just correct, but fast hardware? This paper investigates the question in financial FPGA design, where 5-10 nanoseconds of latency determines competitive advantage and designs iterate continuously as protocols, strategies, and regulations evolve. FinHardBench, a benchmark of 33 financial computing tasks, is presented together with three experiments that mirror the real-world FPGA iteration cycle: generating new modules from specifications, tuning system-level configurations across a 6-stage trading pipeline, and adapting existing modules to specification changes. Evaluation of six LLMs on 1530+ experiment rounds yields three findings: (1) models achieve 19-61% functional correctness with timing degradation up to 13.7$\times$ on specific tasks; (2) in system-level design space exploration, top LLMs converge to the optimal configuration with higher reliability than random search, simulated annealing, and Bayesian optimization baselines (5/5 seeds vs. 0-4/5 at the same 24-round budget); (3) strategy-level specification changes remain unsolved for most models. Across the six models, generation and DSE rankings overlap moderately: the strongest code generator is not the fastest architecture optimizer, and the weakest code generator (MiniMax M2.7) still reaches the system optimum on 4 of 5 seeds. On the tasks in FinHardBench, difficulty tracks training data pattern availability more closely than abstraction level. FinHardBench is released as an open-source benchmark.
Weimin Fu, Hejia Zhang, Minghao Shao et al.· 0 citations
This work presents a comprehensive analysis of contemporary hardware fuzzing techniques applied across three major abstraction layers: Instruction Set Architecture (ISA), microarchitecture, and Register-Transfer Level (RTL). Our study examines key factors including input stimulus quality, mutation strategies, feedback mechanisms, target platforms, reference models, and achieved coverage. We find challenges, goals, and design trade-offs vary significantly across abstraction layers. We further identify several unmet needs in current hardware fuzzing practices, such as intelligent input generation, reliable and scalable golden reference models, expressive feedback channels, and cross-layer integration. Building on these insights, we outline future research directions, including hybrid fuzzing frameworks, AI-assisted test generation, scalable reference models, standardized evaluation metrics and benchmarks, and human-in-the-loop automation for guided exploration and analysis. Together, they aim to unlock efficient, reliable, and comprehensive hardware verification solutions.
Alenkruth Krishnan Murali, Raghul Saravanan, D. SaiManojP et al.· 0 citations
Performance modeling is central to hardware design and software optimization, yet constructing these models requires structured reasoning about computation, data reuse, storage, and movement. We introduce PerfReasoning, a benchmark that evaluates LLMs both as direct performance reasoners and as generators of analytical performance-model code. Given workload, architecture, and mapping specifications, models compare mappings and predict off-chip traffic and buffer requirements. The strongest closed-source models exceed 90% on reasoning-based Q&A, and the best open-weight model reaches 82.4%. However, model construction is substantially harder: while GPT-5.6 Sol exceeds 80% pass rate, all other model configurations average below 15% and vary markedly across runs. Task-specific RL raises a 4B model's mapping-reasoning accuracy by 15.7 points, whereas feedback-free multi-round self-revision prompting is not reliably effective. PerfReasoning exposes the gap between plausible architectural reasoning and reliable performance-model construction. We will publicly release the benchmark to support reproducible evaluation and track future progress.
Da Zhao, K. Sankaralingam, Christos Kozyrakis et al.· 0 citations
A reproducible benchmarking platform that evaluates open-source LLMs on Verilog RTL generation across 50 curated tasks consisting of combinational, sequential, finite state machine (FSM), and mixed designs, enabling reproducible evaluation of generative AI for hardware design workflows.
FPGAgent is the first task-specification-to-executable HLS generation framework experimentally validated on a well-established benchmark, and the value of end-to-end validation is demonstrated.
Tian-Yun Wang, Wenjie Wang, Jian-Guo Yao et al.· 0 citations