Pith. sign in

REVIEW 7 cited by

Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.14074 v1 pith:ZZLEI73G submitted 2025-06-17 cs.LG cs.AR

classification cs.LGcs.AR
keywords designproblemsverificationbenchmarkcvdphardwareagenticcomprehensive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We present the Comprehensive Verilog Design Problems (CVDP) benchmark, a new dataset and infrastructure to advance LLM and agent research in hardware design and verification. CVDP includes 783 problems across 13 task categories, covering RTL generation, verification, debugging, specification alignment, and technical Q&A authored by experienced hardware engineers. Problems are offered in both non-agentic and agentic formats. The benchmark introduces more realistic and challenging contexts than prior work, with state-of-the-art models achieving no more than 34% pass@1 on code generation. Agentic tasks$\unicode{x2013}$especially those involving RTL reuse and verification$\unicode{x2013}$are particularly difficult. Evaluation uses open-source tools and model scoring infrastructure, with comprehension tasks assessed via BLEU and LLM-based judging. CVDP reveals substantial gaps in current model capabilities, underscoring the need for continued research toward robust, real-world hardware design automation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CHIA: An open-source framework for principled, agentic AI-driven hardware/software co-design research

    cs.AR 2026-06 unverdicted novelty 7.0 of 10

    CHIA is an open-source framework for agentic AI-driven hardware/software co-design using CHIA loops as directed cyclic graphs, a tool library, and features for reliable experimentation, shown via five case studies.

  2. Interpretable and Verifiable Hardware Generation with LLM-Driven Stepwise Refinement

    cs.SE 2026-06 unverdicted novelty 7.0 of 10

    Framework uses LLM-driven stepwise application of transformation rules to generate verifiable RTL hardware designs from specifications.

  3. FinHardBench: Can LLMs Generate Latency-Aware Hardware for Financial Computing?

    cs.CL 2026-08 conditional novelty 6.0 of 10

    FinHardBench, 33 financial FPGA tasks, finds LLMs pass functional tests 19-61% of the time and produce many routed-but-wrong designs, while top models tune a 6-stage pipeline to optimal latency more reliably than thre...

  4. Can AI Agents Really Complete RTL-to-GDS? Lessons from Benchmarking Tool-Interactive EDA Workflows

    cs.AI 2026-07 conditional novelty 6.0 of 10

    LLM agents can complete an RTL-to-GDS chip flow, but reliable completion depends on the execution infrastructure, not the foundation model alone.

  5. When LLMs Over-Answer: Measuring and Mitigating Quality Issues in LLM-Based Hardware Description Language Question Answering

    cs.AI 2026-07 conditional novelty 6.0 of 10

    LLM answers to HDL questions are often redundant and verbose; a task-aware multi-agent framework cuts redundancy by 37% and padding by 31% while raising judge-based quality scores.

  6. ChipVerilog: A Large-Scale OpenCores-Derived Benchmark for LLM-Based Verilog RTL Generation

    cs.AR 2026-07 conditional novelty 5.0 of 10

    On 64 large OpenCores-derived Verilog tasks, top LLMs reach 23.6% functional pass@1, 37.5% pass@5, and 0% on designs with two or more submodules, showing hierarchical RTL generation remains unsolved.

  7. Revolution or Hype? Seeking the Limits of Large Models in Hardware Design

    cs.LG 2025-09 conditional novelty 1.0 of 10

    Large models can help early-stage hardware design and verification, but their reliability, data, and precision limits mean traditional EDA algorithms and formal verification remain necessary.

Pith tools