{"id":"3ec6bf9b-46cc-4312-9a5a-0c6d4dac4d1b","arxiv_id":"2509.04905","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":1.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Large models can help early-stage hardware design and verification, but their reliability, data, and precision limits mean traditional EDA algorithms and formal verification remain necessary.","lead":"This paper is a panel summary for ICCAD 2025 in which EDA experts debate whether large AI models will transform chip design or remain hype. It concludes that LLMs and specialized circuit models have real promise but face hard limits in reliability, data scarcity, and precision, and recommends hybrid human-machine flows with formal verification.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The LLM-vs-LCM roadmap rests on an unproven semantic-gap premise, and two key citations appear to point the wrong way.","rationale":"The paper is a panel synthesis and expert-opinion piece, not a formal research claim, so the central qualitative conclusion—large models need formal verification and are not yet ready to replace EDA algorithms—is defensible and well aligned with the current literature. The strongest evidence is the documented commercial deployments (Synopsys assistants, askEDA-style tools) and the consistent expert consensus on reliability limits. The reader's weakest-assumption identification is essentially correct: the paper's LLM/LCM division of labor depends on the claim that text-based LLMs cannot natively capture circuit structure and high-precision numerical reasoning. I partially agree, but I would locate the concern more sharply in the paper's use of evidence. The semantic-gap argument is plausible but categorical; no benchmark is offered to show that current text-based LLMs fail on netlist-level structural tasks. Moreover, the two citations that would ground this claim appear to be misused: [43] is a positive arithmetic-capability result, and [39] is about chain-of-thought data distributions, not EDA assistant value. This does not overturn the paper's qualitative conclusion, but it does mean the 'why LCMs are necessary' argument is weaker than the text suggests. Because the reader already issued CONDITIONAL, my read does not change the verdict; the paper should still be accepted with the conditions of correcting the citation errors and framing the LLM/LCM premise as expert opinion rather than established fact.","tokens_in":13712,"tokens_out":4727,"duration_ms":52900,"concrete_test":"Run a controlled netlist-sensitivity benchmark: take roughly 100 gate-level netlists with one local structural perturbation each; ask a state-of-the-art text-only LLM (with the netlist serialized as text or Verilog, plus few-shot examples) to predict the sign and magnitude of timing or power change, and compare against an LCM/GNN baseline trained on the same data. If the text-only LLM achieves comparable accuracy on non-local effects, the §IV.B/§V.A premise is falsified. In parallel, re-check references [39] and [43] against their abstracts and correct the citations; if [43] indeed reports positive arithmetic results, the numerical-limitation argument loses its cited support.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—LLMs for intent, LCMs for circuit-native optimization—depends on the assertion in §IV.B and §V.A that text-trained LLMs cannot capture circuit structure because a minor netlist change can have non-local effects, and on Markov's claim in §V.A that standard transformers are inefficient at high-precision numerical tasks. This is the load-bearing premise: if a sufficiently capable text-based model can learn graph/structural data, the rationale for separate LCMs and the roadmap weakens. The paper does not demonstrate this premise; it states it. Worse, two supporting citations appear to point the wrong way. Reference [43] is 'Transformers can do arithmetic with the right embeddings'—a positive result showing arithmetic is achievable with appropriate encodings—yet it is cited for the opposite claim that LLMs 'struggle with even basic arithmetic operations.' Reference [39] is a chain-of-thought/data-distribution paper, not a practical EDA-assistant use case, so the Leon Stok paragraph's support for 'LLMs already proving valuable' is misplaced. These are not fatal to the qualitative conclusion, but they show that the claimed fundamental limitation is currently an expert opinion backed by misfiled evidence, not an established result. If that premise is wrong, the paper's sharp LLM/LCM division is less compelling, and the main roadmap becomes one possible design choice rather than a necessity.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a position/panel paper prepared for an ICCAD 2025 panel. It surveys the recent wave of large language models (LLMs) and large circuit models (LCMs) for hardware design, summarizing current tools and benchmarks, enumerating opportunities (RTL generation, verification, design-space exploration, tool orchestration), and cataloguing challenges (reliability/hallucination, semantic gap, data scarcity, explainability). The central thesis is a division of labor: LLMs interpret high-level design intent ('the What') and LCMs, which natively understand circuit structure, perform optimization and implementation ('the How'), with formal verification anchoring all generative outputs. The paper synthesizes opinions from six experts and concludes that large models are a disruptive force but their integration requires solving reliability, data, and precision problems.","tokens_in":14006,"tokens_out":3182,"duration_ms":33960,"significance":"As a synthesis by leading EDA researchers, the paper provides a useful snapshot of the LLM/LCM landscape and articulates a concrete research roadmap. Its strengths are that it names and organizes many recent systems, distinguishes LLMs from LCMs, and explicitly prioritizes verification and trust. The paper is honest about its nature: it contains no new data or derivations, and its value rests on the expertise of the panelists. However, the central roadmap depends on a claim about a 'semantic gap' that is stated rather than demonstrated, and two supporting citations are mischaracterized. If the roadmap is understood as a testable research proposal, the paper is a valuable starting point; if it is taken as an established necessity, the evidence is currently insufficient.","major_comments":[{"comment":"The central LLM/LCM division of labor rests on the premise that general-purpose LLMs trained on text cannot capture circuit structure because 'a minor change in a netlist's structure can trigger serious non-local effects' and 'this causality is not natively captured by models trained on sequential text.' This is asserted, not established. The paper's own survey of LLM-based RTL generation (e.g., RTLCoder, VerilogCoder) shows that text-based models can generate substantially correct RTL, suggesting that the semantic gap is not absolute. As written, the sharp division is one plausible design choice rather than a demonstrated necessity. Please either provide empirical evidence (e.g., controlled comparisons of text-only vs. graph-native models on circuit tasks) or explicitly reframe the roadmap as a hypothesis with falsifiable predictions.","section":"§V.A and §IV.B"},{"comment":"Two citations point the wrong way. (1) Ref. [43] is cited to support 'They struggle with even basic arithmetic operations [43],' but this paper is titled 'Transformers can do arithmetic with the right embeddings' and reports positive results conditional on input embeddings. Citing it as evidence of a fundamental arithmetic limitation misrepresents the source; please replace it with work that actually demonstrates such a limitation, or qualify the claim to reflect that arithmetic accuracy is achievable with suitable encodings. (2) Ref. [39] is cited immediately after 'practical use cases where LLMs are already proving valuable,' but [39] is a study of chain-of-thought reasoning and data distributions, not a collection of EDA use cases. The actual practical examples are cited in [40] and [41]. Either remove [39] or move it to a context where its content is relevant.","section":"§V.A, Refs. [43] and [39]"},{"comment":"Independent of the citation issue, the manuscript treats 'standard transformers are generally inefficient at representing and reasoning over high-precision numerical values' as a settled limitation. This is a load-bearing assumption for the paper's recommendation that LCMs adopt different architectures. It is an expert opinion, not a demonstrated result, and it is important because the entire 'moat' argument for traditional algorithms depends on it. Please distinguish established findings from expert conjecture, and suggest how this claim could be evaluated empirically.","section":"§V.A (Markov's arithmetic claim)"}],"minor_comments":[{"comment":"Typo: 'must use natively handle' should be 'must natively handle'.","section":"§II.B"},{"comment":"Punctuation error: 'with complex and opaque decision-making, .and especially' should be 'with complex and opaque decision-making, and especially'.","section":"§IV.D"},{"comment":"'wholistic' should be 'holistic'. Also 'reveals patterns of PPA optimization gives them' has a subject-verb agreement issue.","section":"§III.C"},{"comment":"Typo in Rolf Drechsler's biography: 'Internationa' should be 'International'.","section":"Panelist Biographies"},{"comment":"Reference [35] has a stray comma in the author list: 'C. K. Jha, ,'.","section":"References [35]"},{"comment":"The timeline figures are dense and have no axes or legend. Consider adding a brief caption explanation of what the colors/positions mean, as this is important for a reader who encounters the figures outside the panel context.","section":"Figures 2 and 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is a panel-position synthesis rather than an empirical study, so the main issues are epistemic framing and citation accuracy. The load-bearing semantic-gap premise needs to be presented as a hypothesis or supported by evidence, and the two misfiled citations should be corrected. The paper also relies substantially on the authors' own prior work (DeepGate, ChatCPU, AutoBench, etc.); this is not disqualifying but should be weighed when assessing independence. With those revisions, the manuscript would be suitable as an authoritative overview."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, quick take: this is a well-organized survey/panel summary, not a research paper. It gives the EDA community a readable map of where LLMs and LCMs stand, and the 'What vs How' split between LLMs (intent) and LCMs (circuit-native optimization) is genuinely clean. If you work in AI-for-EDA, you will find the literature coverage and the panelist positions useful as a shared reference.\n\nWhat it does well: the paper is balanced. It doesn't just cheerlead; the panelists uniformly stress reliability, verification, data scarcity, and sign-off. The conclusion—generative outputs must be anchored in formal verification—is sensible and well-aligned with existing evidence. It also gives due credit to hybrid systems and LLM-plus-tool workflows, which is where most current value is.\n\nSoft spots, in order of real impact. The citation mismatches are real and visible. [39] is a chain-of-thought/data-distribution paper, not support for Stok's practical LLM use cases. [43] is titled 'Transformers can do arithmetic with the right embeddings'—a positive result—yet the text cites it for the claim that LLMs struggle with basic arithmetic. Both are easy to fix but they signal careless sourcing. The load-bearing premise, in IV.B and V.A, is that text-trained LLMs cannot capture circuit structure because a minor netlist change can have non-local effects. That is plausible expert opinion, but it is stated, not demonstrated. If a future LLM or agent handles graph-structured data well, the sharp LLM/LCM division becomes one design choice among several, not a necessity. The paper should acknowledge that as an open question.\n\nThere is also a mild self-referential bias: much of the LCM advocacy is co-authored by Xu, who originated the LCM term. That's not a flaw by itself, but it means the roadmap should be read as advocacy plus expert opinion, not independent evidence.\n\nOverall: for what it is—a foundational text for a panel—it does its job. It has no new data or derivations, so it won't change minds on its own. But as a synthesis it deserves a serious referee: someone should catch the citation errors, and the authors should soften the semantic-gap claim to 'currently underexplored' rather than 'not captured.' I'd accept it for workshop/panel publication after minor revision, and I'd cite it as a useful overview. Bring it to reading group if you want a snapshot of where the field's senior voices stand.","headline":"A useful, balanced panel position paper that reads as expert synthesis rather than research; the LLM/LCM division is a clean framing, but two citations point the wrong way and the central premise is asserted, not tested.","tokens_in":14482,"tokens_out":1866,"would_cite":true,"duration_ms":19453,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that large models in hardware design become trustworthy only when paired with formal verification and a clear split between language models and circuit-native models.","keywords":["large language models","large circuit models","electronic design automation","hardware design","formal verification","RTL generation","PPA optimization","AI reliability"],"falsifier":"A controlled evaluation in which a general-purpose, text-only LLM or generic agentic system matches or beats a circuit-specialized model on held-out gate-level netlists for delay, area, and power would contradict the paper's claim that circuit data needs native LCMs.","tokens_in":13622,"feed_emoji":"🔧","tokens_out":7136,"duration_ms":70232,"temperature":0.7,"pith_summary":"The paper is a position statement, built around a 2025 design-automation panel, asking whether large AI models will transform chip design or remain overhyped. It argues that the honest answer is both: LLMs already help with specification, code snippets, assertions, and debugging, but they hallucinate, cannot be trusted with high-precision numerical work, and are trained on text rather than circuit structure. To move from demonstrations to production, the paper proposes a division of labor — LLMs interpret design intent, circuit-native LCMs handle structural optimization, and formal verification validates every generated artifact. The reader should care because chip bugs can cost millions and cause silicon failures; whether this roadmap works decides if AI-assisted design becomes a routine engineering tool or a permanent niche.","feed_headline":"Formal checks decide if AI chip design is revolution or hype","feed_subtitle":"LLMs read intent, circuit-native models optimize, and only formal checks make the output safe to trust.","key_machinery":"The central mechanism is the 'What versus How' split: LLMs as the front-door natural-language interface ('what' the chip should do), LCMs as the back-end expert reasoning engine ('how' to build it correctly and efficiently), with formal verification as the mandatory trust anchor between them. The paper's motivation rests on the PPA ceiling (heuristics converging to local optima) and the semantic gap (circuit state changes have non-local effects invisible to sequential text models).","core_discovery":"This paper, prepared as the basis for a 2025 design-automation panel debate, argues that large models are neither an instant revolution nor a passing fad. Their real contribution, the authors claim, will come from a division of labor: language models, good at interpreting high-level human intent, should turn specifications into formal descriptions and verification collateral; circuit-native models, trained on logic, topology, and geometry together, should carry the optimization-heavy 'how' work; and formal verification must sit at the center, certifying every generative output before it is trusted. The paper's skeptical corrections — hallucinations, data scarcity, the semantic gap between te","pith_inferences":["If graph-native LLM backbones become practical, the boundary between LLM and LCM may dissolve, with a single model handling both intent and structure; the paper's division of labor is an assumption, not a law.","The paper's formal-verification anchor suggests a testable design: an automated loop where every LLM or LCM edit is checked by equivalence checking or model checking; its cost, not the model, will decide industrial adoption.","Because benchmark saturation is already happening, the next round of progress claims is likely to be on more holistic benchmarks, making independent third-party evaluation of training data the key trust mechanism.","If synthetic data generation matures for circuits, data scarcity may become a shorter-term bottleneck than the semantic gap, reversing the priority of LCM research."],"forward_implications":["LLM-based tools are likely to be adopted early in the design flow for specification translation, testbench and assertion generation, and report triage, but not for final sign-off.","LCMs will need graph-native architectures and multimodal datasets spanning RTL to layout before they can become trusted PPA optimization engines.","Any AI-generated code, assertion, or testbench must pass formal or otherwise deterministic verification before entering production flows.","Benchmarks should shift from RTL code-generation accuracy toward real-world proxy metrics: timing closure, PPA, functional correctness under formal verification, and generalization without data leakage.","Synthetic data generation and privacy-preserving deployment, such as on-premise inference, are necessary to overcome proprietary data scarcity."],"supporting_citations":[{"why":"Surveys existing LLM-for-EDA research, providing the paper's baseline account of how language models entered hardware design.","marker":"[5]"},{"why":"Shows a concrete open-source LLM pipeline for RTL code generation, evidence for the claim that LLMs can draft hardware code.","marker":"[9]"},{"why":"Introduces an agentic Verilog coding system with tool-in-the-loop feedback, used to illustrate current LLM capabilities and limits.","marker":"[10]"},{"why":"Introduces the Large Circuit Model concept and its multimodal circuit-data premise, the paper's key proposed counterweight to text-only LLMs.","marker":"[11]"},{"why":"Frames LLM use in hardware verification and motivates coupling LLMs with formal methods, supporting the paper's central reliability argument.","marker":"[34]"},{"why":"Describes a prompt-verify-repeat loop in hardware verification, direct evidence for the formal-verification-anchored workflow the paper recommends.","marker":"[35]"},{"why":"Demonstrates synthetic tabular data enabling a foundation model for small-data tasks, supporting the paper's call to address EDA data scarcity with synthetic data.","marker":"[42]"},{"why":"Analyzes transformer arithmetic limitations, supporting the claim that standard LLMs are poorly suited to high-precision EDA computations.","marker":"[43]"},{"why":"Reports that reasoning-model output quality degrades as task size grows, cited against the analogy between LLM reasoning and human thinking.","marker":"[44]"}],"fun_headline_variants":["AI chip design: LLMs for intent, circuit models for optimization, formal for trust","For AI hardware, formal verification is the ultimate gatekeeper","Chip design's AI debate: not revolution, not hype—formal verification","LLMs set specs, circuit models optimize, formal checks certify","AI for chips: trusted only after formal verification"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The roadmap assumes that general-purpose language models trained on text cannot natively handle the graph-structured, multimodal nature of circuits, and that separate circuit-native models are therefore necessary.","fun_headline_variants_meta":{"raw":{"variants":["AI chip design: LLMs for intent, circuit models for optimization, formal for trust","For AI hardware, formal verification is the ultimate gatekeeper","Chip design's AI debate: not revolution, not hype—formal verification","LLMs set specs, circuit models optimize, formal checks certify","AI for chips: trusted only after formal verification"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00081,"raw_usage":{"total_tokens":3351,"prompt_tokens":662,"completion_tokens":2689,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":406,"completion_tokens_details":{"reasoning_tokens":2599}},"tokens_in":406,"tokens_out":2689,"duration_ms":18580,"temperature":1.0,"reasoning_tokens":2599,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T05:46:56.471622+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled evaluation in which a general-purpose, text-only LLM or generic agentic system matches or beats a circuit-specialized model on held-out gate-level netlists for delay, area, and power would contradict the paper's claim that circuit data needs native LCMs.","supporting_citations":[{"cited_title":"A survey of research in large language models for electronic design automation,","cited_arxiv_id":null,"evidence_quote":"Surveys existing LLM-for-EDA research, providing the paper's baseline account of how language models entered hardware design."},{"cited_title":"Rtl- coder: Fully open-source and efficient llm-assisted rtl code generation technique,","cited_arxiv_id":null,"evidence_quote":"Shows a concrete open-source LLM pipeline for RTL code generation, evidence for the claim that LLMs can draft hardware code."},{"cited_title":"Verilogcoder: Autonomous Verilog coding agents with graph-based planning and abstract syntax tree (ast)- based waveform tracing tool,","cited_arxiv_id":null,"evidence_quote":"Introduces an agentic Verilog coding system with tool-in-the-loop feedback, used to illustrate current LLM capabilities and limits."},{"cited_title":"Large circuit models: opportunities and challenges,","cited_arxiv_id":null,"evidence_quote":"Introduces the Large Circuit Model concept and its multimodal circuit-data premise, the paper's key proposed counterweight to text-only LLMs."},{"cited_title":"Llms for hardware verification: Frameworks, techniques, and future directions,","cited_arxiv_id":null,"evidence_quote":"Frames LLM use in hardware verification and motivates coupling LLMs with formal methods, supporting the paper's central reliability argument."},{"cited_title":"Prompt. verify. repeat. llms in the hardware verification cycle,","cited_arxiv_id":null,"evidence_quote":"Describes a prompt-verify-repeat loop in hardware verification, direct evidence for the formal-verification-anchored workflow the paper recommends."},{"cited_title":"Accurate predictions on small data with a tabular foundation model,","cited_arxiv_id":null,"evidence_quote":"Demonstrates synthetic tabular data enabling a foundation model for small-data tasks, supporting the paper's call to address EDA data scarcity with synthetic data."},{"cited_title":"Transformers can do arithmetic with the right em- beddings,","cited_arxiv_id":null,"evidence_quote":"Analyzes transformer arithmetic limitations, supporting the claim that standard LLMs are poorly suited to high-precision EDA computations."}],"review_version":1}