Pith. sign in

Paper Citation Record · LEDGER

Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2506.14074.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.14074 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:45:24.830628Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.523834Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4fefa38d-41cd-4acc-9549-bf39075cadd6 · inbound

Revolution or Hype? Seeking the Limits of Large Models in Hardware Design cites this paper.

Revolution or Hype? Seeking the Limits of Large Models in Hardware Design Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T05:50:29.819033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:50:29.819033Z digest=sha256:01586666136b2eb277502279ea0d467bf4e50d601379cd74d1901702ca800c25

Observation b4554590-ea58-42ce-b493-aeecd25816e8 · inbound

HWE-Bench: Benchmarking LLM Agents on Real-World Hardware Bug Repair Tasks cites this paper.

HWE-Bench: Benchmarking LLM Agents on Real-World Hardware Bug Repair Tasks Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:30:18.925789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T11:27:11.257817Z digest=sha256:9bb1e8d7794544a0ba90ed5d3a211d03384d94c74153f30a0748204d96df534d

Observation 6f2649cf-7615-4134-8db7-818d406a1eb7 · inbound

Dr. RTL: Autonomous Agentic RTL Optimization through Tool-Grounded Self-Improvement cites this paper.

Dr. RTL: Autonomous Agentic RTL Optimization through Tool-Grounded Self-Improvement Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:49:56.192583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T10:48:11.646826Z digest=sha256:dc24020bd4394e4d0f3c85d000e81f8c89a47d3a159046c8f21017df3b0bddd8

Observation 05a66429-c6ca-49f4-a6b7-c14695a4b5dc · inbound

Spec2Cov: An Agentic Framework for Code Coverage Closure of Digital Hardware Designs cites this paper.

Spec2Cov: An Agentic Framework for Code Coverage Closure of Digital Hardware Designs Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:17:37.542242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T08:14:18.075112Z digest=sha256:7c04893490beb201f91a28d21e79d23b09a9409a56809a22820382069d1742b2

Observation 3d8b9413-7a79-4d95-a73a-12b631e29fea · inbound

Spec2Cov: An Agentic Framework for Code Coverage Closure of Digital Hardware Designs cites this paper.

Spec2Cov: An Agentic Framework for Code Coverage Closure of Digital Hardware Designs Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:21:24.133526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T10:18:18.439819Z digest=sha256:07627d03bf4ca8b0d4c00e41f50b5204eed7b997db0994428166e339e5d154c7

Observation e9be5ce3-29dd-49ba-a076-528c46bea872 · inbound

Understanding Inference-Time Token Allocation and Coverage Limits in Agentic Hardware Verification cites this paper.

Understanding Inference-Time Token Allocation and Coverage Limits in Agentic Hardware Verification Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:25.037753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T08:00:22.126447Z digest=sha256:0e53456b06a5eadd8758ed8be65bb16d4c4edc978db16924d43172441cb14125

Observation 66379d30-45c0-42e4-963b-2152a4175cc1 · inbound

ChipCraftBrain: Validation-First RTL Generation via Multi-Agent Orchestration cites this paper.

ChipCraftBrain: Validation-First RTL Generation via Multi-Agent Orchestration Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:08.318679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T01:24:08.563327Z digest=sha256:ccc86d11877029df2584681c7a1bbe75af3e191859b1349c104e359b9cce6c7d

Observation c4b63f7b-e5d4-4c85-a536-402b82b2bc90 · inbound

SafeTune: Mitigating Data Poisoning in LLM Fine-Tuning for RTL Code Generation cites this paper.

SafeTune: Mitigating Data Poisoning in LLM Fine-Tuning for RTL Code Generation Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:36:26.586732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T10:16:07.200458Z digest=sha256:2a0d0ed6ca8e427988fbcee8d8cf2c923d19ae07a1cc6c95b9269d34c9966e0d

Observation e755e713-d726-41ca-9ac0-5d652deb0f22 · inbound

RuC: HDL-Agnostic Rule Completion Benchmark Generation cites this paper.

RuC: HDL-Agnostic Rule Completion Benchmark Generation Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:01:29.100776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T08:16:15.514823Z digest=sha256:a0fd5f010958251bd7d7c099a149ab15fc8678bf46e488f1d47ebccbc8aab505

Observation 70ad44aa-d7cd-4b57-b824-2abbb68a642a · inbound

ChipMATE: Multi-Agent Training via Reinforcement Learning for Enhanced RTL Generation cites this paper.

ChipMATE: Multi-Agent Training via Reinforcement Learning for Enhanced RTL Generation Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:05:05.892499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T21:59:20.163621Z digest=sha256:47e3e1a1252cf8b63581a271ea6e576619734b578182adcb01e175da0ba660bd

Observation 98ed2132-5527-46c7-aa76-26bd1d1860bf · inbound

RTL-BenchMT: Dynamic Maintenance of RTL Generation Benchmark Through Agent-Assisted Analysis and Revision cites this paper.

RTL-BenchMT: Dynamic Maintenance of RTL Generation Benchmark Through Agent-Assisted Analysis and Revision Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:42:37.291864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T14:42:16.802186Z digest=sha256:e79ab572add5c34ba33fdbccf39d2dd6a18cfc25f3b56dbd09cd4cd84f24c90e

Observation 1eca5840-96c5-4243-80c5-a90980d25c3c · inbound

AssertLLM2: A Comprehensive LLM Benchmark for Assertion Generation from Design Specifications cites this paper.

AssertLLM2: A Comprehensive LLM Benchmark for Assertion Generation from Design Specifications Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-01T16:25:49.883285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T16:15:54.347613Z digest=sha256:b07cd92b28208dcefd00c42db409c83d5d7986d9ae92237376a2649c9b9d9dd9

Observation 6731d922-2cd0-4f9f-a721-5779b433838a · inbound

CASS-RTL: Correctness-Aware Subspace Steering for RTL Generation with LLMs cites this paper.

CASS-RTL: Correctness-Aware Subspace Steering for RTL Generation with LLMs Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:57:07.337550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T23:06:44.339563Z digest=sha256:b19f6d7e2462fbf17b2b34fc8bfe4d6c963382c46fc633b4b9c4bdfb78786dfd

Observation ad382723-37ab-41e8-a3aa-e1d3cccd64ac · inbound

RTL-BenchLS: A Large-Scale Benchmark for RTL Reasoning and Generation with Large Language Models cites this paper.

RTL-BenchLS: A Large-Scale Benchmark for RTL Reasoning and Generation with Large Language Models Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:47:30.492161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T17:01:24.292776Z digest=sha256:da9989d935a2707650413f33bb7e783b96a4c0ccaf1b6f56f2ad1ae9a5c93df6

Observation 2d10777a-2324-4932-a571-b82ca9fbc2ac · inbound

Structured Testbench Generation for LLM-Driven HDL Design and Verification-Oriented Data Curation cites this paper.

Structured Testbench Generation for LLM-Driven HDL Design and Verification-Oriented Data Curation Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:58:32.897580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T06:50:01.901413Z digest=sha256:c0a9845dab93f5187712658ddd43f67a66c5d4f882f61c43ff0f9d6f75aa1d2c

Observation 7449a3fe-d2c1-4234-9a82-afca4826b75a · inbound

Interpretable and Verifiable Hardware Generation with LLM-Driven Stepwise Refinement cites this paper.

Interpretable and Verifiable Hardware Generation with LLM-Driven Stepwise Refinement Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:28:59.398936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T00:26:54.438203Z digest=sha256:54e27fdef75d689e107d452bf34cb16e2804237014a3b280c3c7ce17342b18ae

Observation 222552d4-52ea-4ec0-abce-269bc2e165ce · inbound

Interpretable and Verifiable Hardware Generation with LLM-Driven Stepwise Refinement cites this paper.

Interpretable and Verifiable Hardware Generation with LLM-Driven Stepwise Refinement Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-15T10:45:43.149976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T10:45:43.149976Z digest=sha256:eea5c028aa2a813eedf237a28d30bda72c5504d164280451c39e944557cc3e94

Observation 90f9680b-adf9-41a8-8e44-a8343be03cc5 · inbound

CHIA: An open-source framework for principled, agentic AI-driven hardware/software co-design research cites this paper.

CHIA: An open-source framework for principled, agentic AI-driven hardware/software co-design research Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.525727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T01:48:58.638094Z digest=sha256:e38b1a2ca86f189e871974f0585d4ab0717367d118538e621361152a59bd9a10

Observation d767d608-007f-4b11-897a-32f4e6770163 · inbound

CHIA: An open-source framework for principled, agentic AI-driven hardware/software co-design research cites this paper.

CHIA: An open-source framework for principled, agentic AI-driven hardware/software co-design research Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:44:39.759592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T10:15:10.471751Z digest=sha256:29bf8b84ba1f4783928bc793d91d00fa49dfb9552ae5edcb0a6cb4ae424ead47

Observation 71f95b0d-fe5c-455d-bc95-6b9a60d62bd0 · inbound

CHIA: An open-source framework for principled, agentic AI-driven hardware/software co-design research cites this paper.

CHIA: An open-source framework for principled, agentic AI-driven hardware/software co-design research Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-02T10:03:40.395732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:03:40.395732Z digest=sha256:09d0459845557c052350bb12819467a579158aadcdc0e8bc0a1046b0899ca0de

Observation 08b2b0cb-1415-4b92-9a54-412c598df5aa · inbound

Agentic Hardware Design as Repository-Level Code Evolution cites this paper.

Agentic Hardware Design as Repository-Level Code Evolution Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-01T18:45:58.335609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T01:49:07.838897Z digest=sha256:eeb19f3ebbdf9ef6f5df412adb5b663d821903cea3bcd503b9d2fcbcec3b6a27

Observation c54ee891-db8a-4e77-9d1d-edc26fd2d2d5 · inbound

ChipVerilog: A Large-Scale OpenCores-Derived Benchmark for LLM-Based Verilog RTL Generation cites this paper.

ChipVerilog: A Large-Scale OpenCores-Derived Benchmark for LLM-Based Verilog RTL Generation Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T07:12:17.834017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:12:17.834017Z digest=sha256:5e1fe14966910c31570bfce9238f7b17b4d815cc8a32b917aacf21a916440f40

Observation c960fe16-40d7-4547-8447-a1ef52a6030a · inbound

When LLMs Over-Answer: Measuring and Mitigating Quality Issues in LLM-Based Hardware Description Language Question Answering cites this paper.

When LLMs Over-Answer: Measuring and Mitigating Quality Issues in LLM-Based Hardware Description Language Question Answering Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T19:12:13.427128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:12:13.427128Z digest=sha256:fa97dadced50a601f9eb420b4f460af81ea2800301d2b7381dc810e0521392d4

Observation 1c791175-4f57-49b6-a500-c8c73ff05263 · inbound

Can AI Agents Really Complete RTL-to-GDS? Lessons from Benchmarking Tool-Interactive EDA Workflows cites this paper.

Can AI Agents Really Complete RTL-to-GDS? Lessons from Benchmarking Tool-Interactive EDA Workflows Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T17:46:04.086561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:46:04.086561Z digest=sha256:7108e2cfc4b4ffcd0cf578984bd7e368431302f2e2347a532f2009eebb77cf12

Observation 60de5fe6-2dfe-447c-b64e-78344d342b27 · inbound

FinHardBench: Can LLMs Generate Latency-Aware Hardware for Financial Computing? cites this paper.

FinHardBench: Can LLMs Generate Latency-Aware Hardware for Financial Computing? Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T00:45:24.830628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:45:24.830628Z digest=sha256:f0b2a8aeebae0abe51b1da25ca9089e9053981e7d3340de836b212c58a5481e4