Skip to content

Can AI Agents Really Complete RTL-to-GDS? Lessons from Benchmarking Tool-Interactive EDA Workflows

Jul 2026 · arXiv.org · Vol abs/2607.17528 · 0 citations · 13 references
Computer Science

TL;DR

The results suggest that robust Agentic EDA requires not only stronger models but also structured tool interfaces, persistent design context, controlled execution, and process-level evaluation.

Abstract

Large language model (LLM) agents are extending electronic design automation (EDA) beyond static RTL generation toward long-horizon, tool-interactive workflows. Yet it remains unclear whether general-purpose coding agents, even with domain-specific EDA skills, can reliably execute an end-to-end RTL-to-GDS flow encompassing synthesis, physical implementation, and engineering change order (ECO) optimization. We evaluate AI agents on a PicoRV32 RTL-to-GDS flow using commercial EDA tools under two timing targets. Their performance is assessed using end-to-end design score, stage completion, and Token ROI, a cost-efficiency metric relating design quality to runtime and cost. Comparing three agent architectures and four foundation models, we derive three practical lessons. First, domain-specific skills improve agents'understanding of individual subtasks but do not ensure reliable completion of a long-horizon EDA flow. Second, agents that achieve similar design progress can still differ by up to 141 times in Token ROI, revealing substantial differences in runtime and cost efficiency. Third, low-level tool-interface mismatches are a major source of physical design failures, particularly when Tcl commands depend on the tool version or execution mode. These results suggest that robust Agentic EDA requires not only stronger models but also structured tool interfaces, persistent design context, controlled execution, and process-level evaluation.

View source

Similar papers

Preprint Aug 2026

StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows

This work systematically study AI products with demonstrated adoption, together with their product workflows and users, to identify real-world tasks for which AI has established practical demand across diverse professional domains and establishes StartupBench as an empirical measure of progress toward E2E completions of real-world user tasks.

Li-Ya Zhu, Xin Ma, Tao Liu et al. · 0 citations
Preprint Aug 2026

DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?

DSAgentBench is introduced, the first benchmark to evaluate whether agents can automate full data-science workflows inside real computer environments, and reveals a substantial capability gap between current agentic systems and real data-science workflows.

Mizanur Rahman, Mohammed Saidul Islam, Ridwan Mahbub et al. · 0 citations
Jul 2026

ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders

ICAE-Bench, a benchmark for evaluating coding agents under interactive project-building settings, starts from a fuzzy product requirement, simulating the dynamic paradigm with an automated User Agent, and introduces three key designs.

Zhongyuan Peng, Dan Huang, Chuyu Zhang et al. · 3 citations · ⚡1
Preprint Aug 2026

Evaluating LLM Trade-offs for Enterprise Automation: Lessons from Workflow Generation in a Production Enterprise Platform

Deploying large language models for AI-driven workflow generation in a production enterprise platform is benchmarked across 29 real-world IT automation scenarios, two generation pipeline architectures, and eight independent runs per prompt-model-pipeline configuration.

Xavier Wrenn, Radoslav Raykov, Aleksandar Angelov et al. · 0 citations
Jul 2026

Beyond Generalist LLMs: Specialist Agentic Systems for Structured Code Workflow Execution

This work introduces a specialist workflow for the transformation of Business Process Model and Notation diagrams into executable agentic workflows and finds that generalist agents generate code inconsistently in both functionality and quality, limiting their suitability for industrial settings where reliability and maintainability are essential.

Harris Borman, Herman Wandabwa, Fu-Sun Yu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.