Pith. sign in

Paper Citation Record · LEDGER

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts

As of 17 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2608.09567.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09567 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:59:21.169998Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact3
  • verified fuzzy11
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a1282627-8509-4760-946d-533861cd7381 · outbound

This paper cites Repeatability in computer systems research.Com- munications of the ACM, 59(3):62–69, 2016.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Repeatability in computer systems research.Com- munications of the ACM, 59(3):62–69, 2016

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:21.741603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:59:21.105922Z digest=sha256:0dfc470ef7f2fc7d8700369a501ce9b5f34285135d37cfe9c158ca6411ecfce8

Observation d1dfbc5e-d755-4bcb-9f3a-b38169a24809 · outbound

This paper cites PentestGPT: Evaluating and harnessing large lan- guage models for automated penetration testing.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts PentestGPT: Evaluating and harnessing large lan- guage models for automated penetration testing

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:21.731735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:59:21.109554Z digest=sha256:074e0afff9622d744c7d7368cfee563ae41f4ff30c67b69cdce95385f795c3d1

Observation 60a5c369-9de9-4bbd-a35c-2fda88132fa2 · outbound

This paper cites PoC-Adapt: Semantic-Aware Automated Vulnerability Reproduction with LLM Multi-Agents and Reinforcement Learning-Driven Adaptive Policy.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts PoC-Adapt: Semantic-Aware Automated Vulnerability Reproduction with LLM Multi-Agents and Reinforcement Learning-Driven Adaptive Policy

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-11T14:59:21.625299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:59:21.112758Z digest=sha256:6bb2b8e11483f9e713744852485a67ab9dce34d2991158665eac8fafad3005a1

Observation c9316492-b658-4e01-b4c2-1351facc0ad4 · outbound

This paper cites Good News for Script Kiddies? Evaluating Large Language Models for Automated Exploit Generation.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Good News for Script Kiddies? Evaluating Large Language Models for Automated Exploit Generation

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-11T14:59:21.609935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:59:21.115774Z digest=sha256:8cb5414f2a38dc03a5e053c0debfd9d17b734d448d0a0640ffdd413b2ef9eedf

Observation 06240eb3-caa5-4ae1-8982-09a13c19c859 · outbound

This paper cites SEC-bench: Automated bench- marking of LLM agents on real-world software security tasks, 2025.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts SEC-bench: Automated bench- marking of LLM agents on real-world software security tasks, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:21.119045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:21.119045Z digest=sha256:f0841293dfa5fdc49eab5f7eab22d30520bbbc23a8bf661988ac035de65a3aaa

Observation 6f542bdc-2dd2-4900-9858-1aab64c12444 · outbound

This paper cites Automated vulnerability validation and verification: A large language model approach, 2025.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Automated vulnerability validation and verification: A large language model approach, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:21.121999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:21.121999Z digest=sha256:c115bad97bc629fb5102f19667ef6950a1e21449e66d5a8bd52890bf15b80982

Observation 5803089d-ea40-4083-a907-0e16f02c80f0 · outbound

This paper cites Shell or nothing: Real-world benchmarks and memory-activated agents for automated penetration testing, 2025.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Shell or nothing: Real-world benchmarks and memory-activated agents for automated penetration testing, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:21.124600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:21.124600Z digest=sha256:91eddd222549cd6baf7be9dada3366149c7ce70a9242814c9ef24e7adfaf5ad9

Observation e7107df3-6e64-43ba-adec-137281b4aaa5 · outbound

This paper cites Cve-2020-1967.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Cve-2020-1967

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:21.721813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:59:21.127846Z digest=sha256:eac4a507c0a069678e4f1f8b22d49ed24e545f9c690da55b71bdb2b48b6d5c14

Observation 7a1bbd3a-9b00-42dc-9851-03062c6b97e7 · outbound

This paper cites Cve-2021-31162.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Cve-2021-31162

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:21.709414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:59:21.130988Z digest=sha256:86d18e79ee15c59259690ef4a9a17d26651a5513eb4c4a44341f4b4e3a618f79

Observation 95fb2193-4302-4002-919a-d4d807e5b642 · outbound

This paper cites Cve-2021-44228.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Cve-2021-44228

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:21.698634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:59:21.133376Z digest=sha256:2b5b8afb715cfdc7b9b3425e5ddb2b4ac473457e8db112f0f72dddbc05ecc0c1

Observation 00495009-4a96-4cb2-93a4-dbefd281da6b · outbound

This paper cites Cve-2022-22816.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Cve-2022-22816

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:21.689690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:59:21.135474Z digest=sha256:051173c88f5fcfbb9260f4ff7489400b0f87f8de3f5f0e19be5e19160ed65e4f

Observation a96a511a-d63a-49de-bbc3-012a39c3ad09 · outbound

This paper cites Cve-2023-0217.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Cve-2023-0217

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:21.679183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:59:21.137405Z digest=sha256:27e8ef8996da5283bd740a6ab2e65ec512677486bfb874ac7ba1beca563b0dd1

Observation 75fdb339-dfdf-40ce-88dc-bcc9626e46a1 · outbound

This paper cites Cve-2023-25676.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Cve-2023-25676

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:21.667701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:59:21.139624Z digest=sha256:5a4520b23d709df08578ea93f396bcc5ebc899bfa5a01a7ade111cd92c4a4171

Observation 051932f2-7da9-4c96-a5b9-452c4a7a54de · outbound

This paper cites Cve-2025-30223.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Cve-2025-30223

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:21.656506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:59:21.142159Z digest=sha256:4b580342dc4fb6c04ce5eb11d3daeb48b9c3b6c5e947b95224ba987361e10048

Observation 121e8746-34d5-4006-b594-5d76ec9184cd · outbound

This paper cites FaultLine: Automated Proof-of-Vulnerability Generation Using LLM Agents.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts FaultLine: Automated Proof-of-Vulnerability Generation Using LLM Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:21.144083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:21.144083Z digest=sha256:70f860ff61e2a112dd9c407d67ca2a3e8194595cfb2ac0a4433f8a4a87dbb550

Observation cc4577a1-a9f2-45d0-83c6-dd8b8b9e68ce · outbound

This paper cites ”Get in Researchers; We’re Measuring Reproducibility”: A reproducibility study of machine learning papers in tier 1 se- curity conferences.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts ”Get in Researchers; We’re Measuring Reproducibility”: A reproducibility study of machine learning papers in tier 1 se- curity conferences

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:21.146899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:21.146899Z digest=sha256:9846ca79eeda991521a297ecec8154f44c6e49d4c433eb7cfa203185f6b081fb

Observation 0e329e66-dde9-4e57-8ad4-3a7fb2764e8c · outbound

This paper cites Artificial Intelligence as the New Hacker: Developing Agents for Offensive Security.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Artificial Intelligence as the New Hacker: Developing Agents for Offensive Security

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-11T14:59:21.381995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:59:21.150018Z digest=sha256:0acae1a44688d56117e606e5453b15dd85cfa7a047faa1d384fe5c771b77da67

Observation d484c919-ee1f-4bdb-8d38-5aea5222ff80 · outbound

This paper cites Contract- Tinker: LLM-empowered vulnerability repair for real-world smart contracts.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Contract- Tinker: LLM-empowered vulnerability repair for real-world smart contracts

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:21.646448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:59:21.153315Z digest=sha256:485b85321b14495f48c565bdaf824079edb6d77ed9df07dcbd72f06ae84361b0

Observation ebc02d0a-a208-4ec6-8639-f29b534e7d7c · outbound

This paper cites PATCHEV AL: A new benchmark for evaluating LLMs on patching real-world vulnerabilities, 2025.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts PATCHEV AL: A new benchmark for evaluating LLMs on patching real-world vulnerabilities, 2025

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:21.156187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:21.156187Z digest=sha256:b12323a8a1bb40bd1855cde9cb37a38ae408b2edf3d37ab050d691904dc56db9

Observation 9ece5547-9224-41f4-aff0-b4fa6c490bc6 · outbound

This paper cites AutoPT: How Far Are We from the End2End Automated Web Penetration Testing?.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts AutoPT: How Far Are We from the End2End Automated Web Penetration Testing?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:21.159703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:21.159703Z digest=sha256:2b04c6425f0c0947bafeff4cf408eed7f0f1a616dd3249638f1851b4d175d732

Observation 3c4d597c-d324-4f3c-a450-2a8e05b64f1a · outbound

This paper cites Patch validation in automated vulnerability repair, 2026.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Patch validation in automated vulnerability repair, 2026

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:21.164056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:21.164056Z digest=sha256:dfa10a429e60a739eadb675150e7d1978d46cf70e81d8f4cae120305539519d0

Observation 43af46b6-aeb0-426a-878b-a0fd40de25c3 · outbound

This paper cites Zhang, Neil Perry, Riya Dulepet, Joey Ji, Celeste Menders, Justin W.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts Zhang, Neil Perry, Riya Dulepet, Joey Ji, Celeste Menders, Justin W

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:21.636028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:59:21.166818Z digest=sha256:e98cac33a6854662c9b37533771fb11389333f319c83913f6244695a12002dc7

Observation 93a5464a-3e34-4fb2-80e2-97f2de3efcc9 · outbound

This paper cites CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:21.169998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:21.169998Z digest=sha256:d3efb4e70508f361e9d82b54d5fb4e04050c90777c87bb00ba84dc4983af25a6

Pith citing papers

No inbound Pith citation observations are available.