{"id":"0d5b9df2-c1c6-4632-9ee9-0cafe3a7d912","arxiv_id":"2412.09745","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":6,"one_line_summary":"AiEDA combines LLM agents with open-source EDA tools in a four-stage concept-to-GDSII flow, but it has not yet demonstrated the full flow on its KWS case study.","lead":"AiEDA is a proposed software framework that uses AI agents and open-source chip design tools to move a digital circuit from a natural language specification to a final chip layout. The paper's evidence is preliminary: one small FIFO was taken through layout, and a keyword-spotting design was explored at architecture level but not finished end to end.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The end-to-end concept-to-GDSII claim rests on an unreported one-sentence FIFO run; the KWS flow is explicitly unfinished, and agentic control over physical design is left to future work.","rationale":"The reader's rejection is justified. I read the paper in good faith: it is clearly labeled preliminary, and the authors are transparent that work is ongoing. However, the abstract's wording 'demonstrated through the design of an ultra-low-power digital ASIC for KWS' goes beyond the evidence. The only executed flow is a trivial FIFO, and even that lacks all supporting artifacts. More importantly, the framework as specified does not fulfill the 'agentic end-to-end' claim: physical design is executed by OpenROAD/Magic without agentic feedback loops, which the paper explicitly defers to future work. So the load-bearing assumption that an LLM agent can close the loop to GDSII is unsupported. A single reproducible FIFO run with logs would materially improve the submission; a completed KWS GDSII would be even stronger. The verdict should remain rejection, not because the idea is impossible, but because the central claim is not evidenced. I found no basis to soften the rejection; the missing artifacts and unfinished demonstration are exactly what a high-confidence reject requires.","tokens_in":6487,"tokens_out":2816,"duration_ms":28743,"concrete_test":"Obtain or reconstruct the Section V.A FIFO experiment: run the described LangGraph/GPT-4o flow against the Sky130 PDK and publish all artifacts from the run, including the final GDSII, a DRC-clean report, an OpenSTA timing report with slack, and the exact prompt/reflection transcript. Independently run the same FIFO specification through a conventional non-agentic Sky130 flow (e.g., standard OpenROAD) and compare area and timing; if the agentic run does not produce a DRC-clean GDSII, or if its quality is not reported, the core demonstration has not been made.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of the abstract is that AiEDA is 'demonstrated through the design of an ultra-low-power digital ASIC for KeyWord Spotting,' taking a design from conceptual specification to GDSII. The body does not support this. Section V.A reports a 6-bit, 32-depth FIFO run through RTL, synthesis, and netlist stages to GDSII in one sentence, with no RTL listing, synthesis script, timing report, DRC/LVS output, or area/power numbers. Section V.C states the full KWS end-to-end design is 'still in progress,' so the 'demonstrated' KWS flow does not exist. Moreover, the framework described in Section III does not actually apply agentic AI to physical design: it says OpenROAD provides feedback mechanisms and 'Integrating agentic AI into these feedback loops is a potential area for future exploration.' Thus the load-bearing premise that LLM agents can steer the open-source EDA loop from RTL to clean, fabricatable GDSII is assumed, not demonstrated. The FIFO result could be evidence for this premise if run artifacts were provided; without them it is an unverifiable assertion.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes AiEDA, an agentic AI framework intended to guide digital ASIC design from a natural-language specification through architecture exploration, RTL generation, synthesis, and physical design to GDSII using open-source tools. The framework is organized into four phases: architecture design (LLM-driven Python modeling and reflection), RTL design (LLM-generated Verilog with Icarus simulation feedback), netlist synthesis (Yosys and OpenSTA with LLM-assisted timing fixes), and physical design (OpenROAD and Magic). The authors present a keyword spotting (KWS) case study in which an LLM selects a 4 kHz audio bandwidth, 7-bit fixed-point precision, a 32-point FFT, shift-based coefficients, and rectangular Mel filters. The only reported RTL-to-GDSII result is a 6-bit, 32-depth FIFO, described in one sentence. The paper states that the full KWS end-to-end flow is still in progress.","tokens_in":6721,"tokens_out":6115,"duration_ms":58573,"significance":"If validated, AiEDA would extend LLM-based hardware design beyond HDL generation toward an integrated open-source design flow, and the KWS architecture exploration is a plausible demonstration vehicle. The paper provides a clear survey of related work (DAVE, Verigen, ChipNeMo, MG-Verilog, GPT4AIChip, AutoChip) and a systematic presentation of the proposed framework. However, the contribution is currently a proposal: the central demonstration is incomplete, the single GDSII run lacks all supporting artifacts, and the physical-design phase is explicitly not yet agentic. No code, logs, timing reports, or quantitative measurements are provided, so the reported area/power benefits cannot be assessed. The paper's honest disclosure of its status does not compensate for the absence of evidence for the abstract's claim of a demonstrated end-to-end flow.","major_comments":[{"comment":"The abstract states that AiEDA is 'demonstrated through the design of an ultra-low-power digital ASIC for KeyWord Spotting,' but Section V.C says the full end-to-end KWS design is 'still in progress' and that the goal is to complete it before the conference date. No GDSII result, netlist, timing report, or power report for the KWS design appears anywhere in the manuscript. The central claim of a demonstrated concept-to-GDSII flow is therefore unsupported by the body of the paper.","section":"Abstract and Section V.C"},{"comment":"The only RTL-to-GDSII validation is a single sentence describing a 6-bit, 32-depth FIFO: 'The design process began with a prompt specifying the FIFO's requirements and concluded with the generation of the GDSII layout.' No RTL listing, synthesis script, OpenSTA timing report, DRC/LVS result, or area/power number is provided. Without these artifacts, the reader cannot verify that the flow completed successfully, that the GDSII is clean, or that agentic LLM control, rather than manual intervention, produced the result.","section":"Section V.A"},{"comment":"The Physical design phase explicitly states: 'Integrating agentic AI into these feedback loops is a potential area for future exploration.' Thus the current framework does not apply agentic AI to placement, clock tree synthesis, routing, or DRC closure. As a result, the paper's framing of AiEDA as an agentic flow that 'streamline[s] the transition from conceptual design to GDSII layout' overstates what is implemented; at most, the agentic portion covers architecture, RTL, and synthesis, with backend physical design performed by conventional OpenROAD/Magic tool flows.","section":"Section III, Physical design"},{"comment":"The quantitative architectural claims are not supported by any simulation or measurement. The paper asserts that a 4 kHz bandwidth 'will lead to an approximately 5x reduction in power consumption' and that 7-bit precision 'will yield around a 2x reduction in both area and power consumption,' but no power model, synthesis results, baseline comparison, or measurement is presented. Likewise, the statements that spectral leakage is 'limited to 10%' and that a 32-point FFT keeps accuracy loss 'within 25%' are given without plots, scripts, or evaluation data. These numbers may be reasonable estimates, but they are not demonstrated results.","section":"Section V.B"},{"comment":"The RTL design loop is described as repeated until 'functional verification is successful,' but the test benches are also generated by the LLM, and no independent verification, coverage results, or assertion-based checks are reported. The framework's reliance on LLM self-reflection as the sole correctness mechanism is a correctness-risk concern; a concrete test would be to provide a separately authored testbench or formal property suite, neither of which appears in the manuscript.","section":"Section III, RTL design"}],"minor_comments":[{"comment":"The phrase 'refer to Figure Section III' is ambiguous; it should refer to a specific figure or section, e.g., Section III or Figure 2.","section":"Section V.A"},{"comment":"'Reinforcement Learining with Human Feedback' contains a typo; it should be 'Learning'.","section":"Section VI"},{"comment":"The URL in reference [3] contains a duplicated and mangled segment ('agentic-wagentic...orkflows'); it should be a single clean URL.","section":"Reference [3]"},{"comment":"The phrase 'design of a ultra low power digital ASIC' should be 'design of an ultra-low-power digital ASIC' for grammatical correctness.","section":"Introduction"},{"comment":"The authors state that the design tools are 'still in active development' and that the project will be released as open-source once stable; adding a public repository or appendix with the FIFO scripts would strengthen reproducibility.","section":"Section V"}],"recommendation":"reject","confidential_remarks":"The manuscript is more of a position/proposal than a completed research contribution. The authors are transparent about the ongoing status, but the submission does not meet the bar for an archival paper; a demonstration with artifacts and independent verification is needed. A future submission could be appropriate once the KWS flow is completed and the physical-design phase is actually agentic."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. First, it is a genuine proposal, not a fake one: the authors describe a four-stage agentic flow (architecture, RTL, synthesis, physical design) that integrates LLM feedback loops with open-source tools, and they are upfront in the body that the work is ongoing. Second, the abstract oversells it. It says the framework is 'demonstrated' through an ultra-low-power KWS ASIC, but Section V.C admits the full end-to-end KWS design is still in progress, and the only complete GDS run is a 6-bit FIFO described in one sentence with no logs, no timing reports, no DRC output, no area or power numbers. What is actually new here is the integration. Prior work like AutoChip and GPT4AIChip stops at Verilog generation or FPGA HLS; AiEDA sketches a path from natural-language specification through synthesis, timing repair, and OpenROAD-based physical design to GDSII. The related work section is fair, and the authors correctly flag that bringing agentic AI into the physical-design feedback loops is future work. That honesty counts for something. The soft spots are real, and they are load-bearing. The claimed 5x power reduction from lowering bandwidth to 4 kHz and 2x area/power reduction from 7-bit precision are not backed by any simulation, measurement, power model, or baseline. The numbers may be plausible, but they are asserted. Similarly, the FIFO-to-GDS validation is unverifiable as reported. The LLM also serves as both generator and evaluator of design choices, which introduces a circularity that the paper does not address. None of this means the framework is impossible; it means the submitted evidence does not support the abstract's central claim. Who is this for? Researchers working on LLM-based EDA and anyone interested in agentic design flows. It reads like a workshop paper or a position statement. I would not cite it as a validated system, but I might bring it to a reading group to discuss the architecture and open questions. My recommendation: do not desk-reject out of hand - send it to referees, because the framework idea is important enough and plausible enough to merit careful review. But expect the reviewers to send it back for major revision, and tell the authors to release code, provide the FIFO artifacts, and complete the KWS run before claiming a demonstration.","headline":"A clear, honest proposal for an agentic LLM-driven ASIC flow, but the abstract claims a demonstration the body does not provide: one unverifiable FIFO GDS run and an unfinished KWS design.","tokens_in":746,"tokens_out":984,"would_cite":false,"duration_ms":24590,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AiEDA proposes an agentic AI design flow that takes a digital ASIC from a natural-language specification to GDSII layout using open-source EDA tools.","keywords":["agentic AI workflow","digital ASIC design","large language models","HDL generation","open-source EDA","GDSII layout","keyword spotting","retrieval-augmented generation"],"falsifier":"Take a non-trivial component, such as the keyword-spotting front-end, and run AiEDA end-to-end from prompt to GDSII without human RTL edits; then check whether the result passes design-rule and layout-vs-schematic checks, meets timing at the target frequency, and matches the architectural Python model. A run that requires expert manual fixes to close timing or clear DRC errors would show the loop has not closed autonomously.","tokens_in":6258,"feed_emoji":"⚙️","tokens_out":7778,"duration_ms":73395,"temperature":0.7,"pith_summary":"The paper proposes AiEDA, an agentic design framework in which LLM agents coordinate open-source EDA tools across four stages—architecture, RTL, synthesis, and physical design—to take a digital ASIC from a natural-language specification to a GDSII layout ready for fabrication. If it works, chip designers would describe intent in words and let agents iterate with tool feedback until verification, timing, and design-rule checks pass, promising large productivity gains. The demonstration is preliminary: a simple 6-bit FIFO was taken from RTL to GDSII, and architectural exploration of an ultra-low-power keyword-spotting ASIC produced bandwidth and precision choices. The end-to-end KWS design is stated as ongoing work rather than completed.","feed_headline":"AiEDA lets LLM agents steer chips from spec to GDSII","feed_subtitle":"Open-source EDA tools plus self-reflecting agents target a keyword-spotting ASIC, though the full flow is still in progress.","key_machinery":"The load-bearing mechanism is the per-stage feedback loop between an LLM and an EDA tool: a design prompt or reflection prompt is given to the LLM, the LLM produces code or commands, the tool runs, and the tool output is fed back to the LLM for analysis and correction until the stage succeeds. The framework implements this loop in four stages: architecture (Python models and analysis scripts), RTL (Verilog generation with simulator feedback), synthesis (netlist generation with static timing feedback and LLM-generated corrective actions), and physical design (placement, routing, and GDSII generation). Retrieval-augmented generation supplies the LLM with relevant Verilog examples, and the designer can intervene at any point. The same loop is what would turn a specification into a layout without day-to-day expert involvement.","core_discovery":"On the paper's own terms, the central claim is that an agentic design flow—an LLM that plans, invokes EDA tools, reads their feedback, and corrects itself in a loop—can carry a digital ASIC from a natural-language specification through architecture modeling, RTL generation and verification, synthesis and timing closure, and physical design to a GDSII layout. The authors demonstrate this with two parts: a completed but briefly reported run of a 6-bit, 32-depth FIFO from RTL to GDSII using the Sky130 process, and an architecture-stage exploration of an ultra-low-power keyword-spotting ASIC in which the LLM chose a 4 kHz bandwidth, 7-bit precision, a 32-point FFT, and shift-and-add coefficients. The paper is explicit that the full end-to-end KWS design is still in progress, so the claim about a fully autonomous flow is a proposal with preliminary support rather than a finished demonstration.","pith_inferences":["Inference: The practical test of AiEDA is whether LLM-generated timing fixes generalize; a useful benchmark would compare agent corrections against expert fixes on a set of failing timing paths from real designs.","Inference: The framework's dependence on tool feedback means its ceiling is set by the quality of that feedback; structured summarization of long tool logs may be needed before LLM reasoning scales to designs with thousands of violations.","Inference: If the KWS design completes, the same agentic structure could extend beyond digital ASICs to verification and mixed-signal tasks, where feedback loops are less crisp and the opportunity for automation is large.","Inference: Because the completed FIFO run is reported without logs, timing reports, or DRC results, the GDSII claim should be read as a planned capability rather than a measured one."],"forward_implications":["A designer could start from a paragraph specification and obtain synthesizable, verified RTL plus a GDSII layout through an open-source toolchain, removing much of the tool-integration burden.","The self-correction loop could absorb routine simulation, timing, and design-rule fixes that currently consume expert time, letting designers focus on architectural trade-offs.","The architecture-stage loop can surface power, area, and accuracy trade-offs early, as it did for the KWS design's bandwidth and bit-width choices.","If the full flow completes, the same four-stage structure could serve as a template for other digital blocks, with each stage's tool feedback as the only requirement."],"supporting_citations":[{"why":"Supplies the open-source backend tool suite that performs placement, clock-tree synthesis, routing, and optimization in the physical-design stage.","marker":"[2]"},{"why":"Provides the open-source digital-flow context that the framework builds on for backend tasks.","marker":"[11]"},{"why":"Early demonstration that an LLM can derive Verilog from English, which motivates the RTL-generation stage.","marker":"[5]"},{"why":"Shows a fine-tuned LLM generating Verilog, supporting the use of LLMs for HDL generation.","marker":"[6]"},{"why":"Prior agentic loop with LLM-generated hardware and tool feedback, which AiEDA extends to the ASIC flow.","marker":"[9]"},{"why":"Closest predecessor for the iterative LLM-plus-simulation loop used in the RTL stage.","marker":"[10]"},{"why":"Supplies the keyword-spotting architecture and ultra-low-power MFCC engine used as the case study.","marker":"[12]"},{"why":"Provides the Verilog dataset for retrieval-augmented generation that enriches the LLM's context.","marker":"[8]"},{"why":"Provides the agent graph framework used to build the multi-stage agentic flow.","marker":"[15]"}],"fun_headline_variants":["Agentic AI flow steers ASIC from spec to GDSII","LLM agents run EDA tools to carry chips to GDSII","AiEDA: AI agents automate digital ASIC design steps","Self-correcting agents target spec-to-layout chip design"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"That an LLM can keep driving the open-source EDA loop to clean, fabricatable results using only tool feedback and reflection prompts, with the designer stepping in occasionally rather than as a daily expert.","fun_headline_variants_meta":{"raw":{"variants":["Agentic AI flow steers ASIC from spec to GDSII","LLM agents run EDA tools to carry chips to GDSII","AiEDA: AI agents automate digital ASIC design steps","Self-correcting agents target spec-to-layout chip design"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000503,"raw_usage":{"total_tokens":2442,"prompt_tokens":913,"completion_tokens":1529,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":529,"completion_tokens_details":{"reasoning_tokens":1455}},"tokens_in":529,"tokens_out":1529,"duration_ms":15412,"temperature":1.0,"reasoning_tokens":1455,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T16:46:16.471074+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a non-trivial component, such as the keyword-spotting front-end, and run AiEDA end-to-end from prompt to GDSII without human RTL edits; then check whether the result passes design-rule and layout-vs-schematic checks, meets timing at the target frequency, and matches the architectural Python model. A run that requires expert manual fixes to close timing or clear DRC errors would show the loop has not closed autonomously.","supporting_citations":[{"cited_title":"Openroad: Toward a self-driving, open-source digital layout implementation tool chain,","cited_arxiv_id":null,"evidence_quote":"Supplies the open-source backend tool suite that performs placement, clock-tree synthesis, routing, and optimization in the physical-design stage."},{"cited_title":"Toward an open-source digital flow: First learnings from the openroad project,","cited_arxiv_id":null,"evidence_quote":"Provides the open-source digital-flow context that the framework builds on for backend tasks."},{"cited_title":"Dave: Deriving automatically verilog from english,","cited_arxiv_id":null,"evidence_quote":"Early demonstration that an LLM can derive Verilog from English, which motivates the RTL-generation stage."},{"cited_title":"0.08mm2 128nw mfcc engine for ultra-low power, always-on smart sensing applications,","cited_arxiv_id":null,"evidence_quote":"Supplies the keyword-spotting architecture and ultra-low-power MFCC engine used as the case study."},{"cited_title":"Langgraph: Build resilient language agents as graphs,","cited_arxiv_id":null,"evidence_quote":"Provides the agent graph framework used to build the multi-stage agentic flow."}],"review_version":1}