Pith. sign in

Paper Citation Record · LEDGER

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts

As of 17 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2608.09567.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09567 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:59:21.169998Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact3
  • verified fuzzy11
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a1282627-8509-4760-946d-533861cd7381 · outbound

This paper cites Repeatability in computer systems research.Com- munications of the ACM, 59(3):62–69, 2016.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Repeatability in computer systems research.Com- munications of the ACM, 59(3):62–69, 2016

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:21.741603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:59:21.105922Z digest=sha256:4535f167acbd8288604cb20dceb152c52c77e38e9164ea406d26ada918bf822f

Observation d1dfbc5e-d755-4bcb-9f3a-b38169a24809 · outbound

This paper cites PentestGPT: Evaluating and harnessing large lan- guage models for automated penetration testing.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts PentestGPT: Evaluating and harnessing large lan- guage models for automated penetration testing

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:21.731735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:59:21.109554Z digest=sha256:c0aa4bcc634d2a63d0f89bf5371fcc5a38bd460b3fd6a743ec18c9e1fa283f53

Observation 60a5c369-9de9-4bbd-a35c-2fda88132fa2 · outbound

This paper cites PoC-Adapt: Semantic-Aware Automated Vulnerability Reproduction with LLM Multi-Agents and Reinforcement Learning-Driven Adaptive Policy.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts PoC-Adapt: Semantic-Aware Automated Vulnerability Reproduction with LLM Multi-Agents and Reinforcement Learning-Driven Adaptive Policy

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-11T14:59:21.625299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:59:21.112758Z digest=sha256:ed28f2e388e9ed8e1904f300967393db12af5ea86a982fd9724dc95d5c0f0f6b

Observation c9316492-b658-4e01-b4c2-1351facc0ad4 · outbound

This paper cites Good News for Script Kiddies? Evaluating Large Language Models for Automated Exploit Generation.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Good News for Script Kiddies? Evaluating Large Language Models for Automated Exploit Generation

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-11T14:59:21.609935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:59:21.115774Z digest=sha256:6e7fafdb37c4fecd35b0f690b6e4b585c0c5088556980a6dd5a9902585f637ed

Observation 06240eb3-caa5-4ae1-8982-09a13c19c859 · outbound

This paper cites SEC-bench: Automated bench- marking of LLM agents on real-world software security tasks, 2025.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts SEC-bench: Automated bench- marking of LLM agents on real-world software security tasks, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:21.119045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:21.119045Z digest=sha256:ad7640f45c323cce66cb10ca5990cdae0f88f8dec219ae52c81feed15594713b

Observation 6f542bdc-2dd2-4900-9858-1aab64c12444 · outbound

This paper cites Automated vulnerability validation and verification: A large language model approach, 2025.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Automated vulnerability validation and verification: A large language model approach, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:21.121999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:21.121999Z digest=sha256:f434905d4a1706273009fb3f0e12660ab887161825722254a1dfa1a0cbdc660c

Observation 5803089d-ea40-4083-a907-0e16f02c80f0 · outbound

This paper cites Shell or nothing: Real-world benchmarks and memory-activated agents for automated penetration testing, 2025.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Shell or nothing: Real-world benchmarks and memory-activated agents for automated penetration testing, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:21.124600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:21.124600Z digest=sha256:cd64edfe5223636da2957371a4993cc86729d79a7cd59670b6d7c1c273961024

Observation e7107df3-6e64-43ba-adec-137281b4aaa5 · outbound

This paper cites Cve-2020-1967.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Cve-2020-1967

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:21.721813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:59:21.127846Z digest=sha256:4d160d94e1d053525f11d78fdd7cba15dab4241f5ff4a2b3a007bd543fd933e2

Observation 7a1bbd3a-9b00-42dc-9851-03062c6b97e7 · outbound

This paper cites Cve-2021-31162.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Cve-2021-31162

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:21.709414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:59:21.130988Z digest=sha256:a11034cc5ef0f96b8e1a338bb559365a84339c12744ca0dc4997d0c176dd4f54

Observation 95fb2193-4302-4002-919a-d4d807e5b642 · outbound

This paper cites Cve-2021-44228.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Cve-2021-44228

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:21.698634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:59:21.133376Z digest=sha256:8aacca996e13ede51f99eeda5d906d794a705e4401f131db44aa3e25aa14c14a

Observation 00495009-4a96-4cb2-93a4-dbefd281da6b · outbound

This paper cites Cve-2022-22816.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Cve-2022-22816

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:21.689690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:59:21.135474Z digest=sha256:37649d8f6364727dba6d056d08a427d38250f7ae18ff791a6b995ccc147f039f

Observation a96a511a-d63a-49de-bbc3-012a39c3ad09 · outbound

This paper cites Cve-2023-0217.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Cve-2023-0217

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:21.679183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:59:21.137405Z digest=sha256:8a219f63d6ee89fe2186fb71b9e01e331b582f6286a9f2d7e63cc6b7aaaf14e3

Observation 75fdb339-dfdf-40ce-88dc-bcc9626e46a1 · outbound

This paper cites Cve-2023-25676.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Cve-2023-25676

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:21.667701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:59:21.139624Z digest=sha256:e39e2f805984c8fc51710d99d0403518858bffa581ed08bde25267103b0744eb

Observation 051932f2-7da9-4c96-a5b9-452c4a7a54de · outbound

This paper cites Cve-2025-30223.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Cve-2025-30223

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:21.656506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:59:21.142159Z digest=sha256:25bfc62d3a856534dc0c6b3e844a7fa383b0d60c0acce867514b07e68770d445

Observation 121e8746-34d5-4006-b594-5d76ec9184cd · outbound

This paper cites FaultLine: Automated Proof-of-Vulnerability Generation Using LLM Agents.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts FaultLine: Automated Proof-of-Vulnerability Generation Using LLM Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:21.144083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:21.144083Z digest=sha256:74207725439b64c3f64485be1e6f6f27c36406417122612bc424f819b234c42c

Observation cc4577a1-a9f2-45d0-83c6-dd8b8b9e68ce · outbound

This paper cites ”Get in Researchers; We’re Measuring Reproducibility”: A reproducibility study of machine learning papers in tier 1 se- curity conferences.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts ”Get in Researchers; We’re Measuring Reproducibility”: A reproducibility study of machine learning papers in tier 1 se- curity conferences

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:21.146899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:21.146899Z digest=sha256:4d9369af45ab64df1355b45ce0213532f4bb16b020057cc81405ce3e8f813581

Observation 0e329e66-dde9-4e57-8ad4-3a7fb2764e8c · outbound

This paper cites Artificial Intelligence as the New Hacker: Developing Agents for Offensive Security.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Artificial Intelligence as the New Hacker: Developing Agents for Offensive Security

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-11T14:59:21.381995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:59:21.150018Z digest=sha256:714712eeef7a9036479b109820c191a376430359d6a96da59b8b2721d6d6fd1b

Observation d484c919-ee1f-4bdb-8d38-5aea5222ff80 · outbound

This paper cites Contract- Tinker: LLM-empowered vulnerability repair for real-world smart contracts.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Contract- Tinker: LLM-empowered vulnerability repair for real-world smart contracts

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:21.646448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:59:21.153315Z digest=sha256:4f09ef3284932ae9d5376aac6fb8ae0a35ed1c245a8bc859334f7b25b216a6a1

Observation ebc02d0a-a208-4ec6-8639-f29b534e7d7c · outbound

This paper cites PATCHEV AL: A new benchmark for evaluating LLMs on patching real-world vulnerabilities, 2025.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts PATCHEV AL: A new benchmark for evaluating LLMs on patching real-world vulnerabilities, 2025

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:21.156187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:21.156187Z digest=sha256:cec3e92dada6a44812596881d5bcaf4a4336b9941b198ac1080a5558ec198655

Observation 9ece5547-9224-41f4-aff0-b4fa6c490bc6 · outbound

This paper cites AutoPT: How Far Are We from the End2End Automated Web Penetration Testing?.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts AutoPT: How Far Are We from the End2End Automated Web Penetration Testing?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:21.159703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:21.159703Z digest=sha256:6712d2397826c4a61097c32799604b69a242a5b20b01d127e5daef2304cbb91b

Observation 3c4d597c-d324-4f3c-a450-2a8e05b64f1a · outbound

This paper cites Patch validation in automated vulnerability repair, 2026.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Patch validation in automated vulnerability repair, 2026

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:21.164056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:21.164056Z digest=sha256:aa41fb42ac5b607b235a308ac967ea1589d7397444ffda54c6e7a9668ae845f3

Observation 43af46b6-aeb0-426a-878b-a0fd40de25c3 · outbound

This paper cites Zhang, Neil Perry, Riya Dulepet, Joey Ji, Celeste Menders, Justin W.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Zhang, Neil Perry, Riya Dulepet, Joey Ji, Celeste Menders, Justin W

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:21.636028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:59:21.166818Z digest=sha256:318c925f512c99a980642bdbcb47524690ac9d6ee83d21fcdaa3c110a9b89368

Observation 93a5464a-3e34-4fb2-80e2-97f2de3efcc9 · outbound

This paper cites CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:21.169998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:21.169998Z digest=sha256:5dde8895ee675be62a055cea5ee36c3a5c6f829e3d791ea075251f8d73a7d8ff

Pith citing papers

No inbound Pith citation observations are available.