Pith. sign in

Paper Citation Record · LEDGER

CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2503.17332.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.17332 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:16:52.653903Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T20:16:29.518680Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 052abd75-c0a6-49bb-93a5-0c77cb9ba8c9 · inbound

LLMs unlock new paths to monetizing exploits cites this paper.

LLMs unlock new paths to monetizing exploits CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T20:58:10.500385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:58:10.500385Z digest=sha256:00407dbce91eaa7307f394c8823ad0ab41ae0b9f702c770116104329fb9ec1aa

Observation 35e5a63f-ecca-4391-9498-c99b2046f1f8 · inbound

Eradicating the Unseen: Detecting, Exploiting, and Remediating a Path Traversal Vulnerability across GitHub cites this paper.

Eradicating the Unseen: Detecting, Exploiting, and Remediating a Path Traversal Vulnerability across GitHub CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 139

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:11.023893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:11.023893Z digest=sha256:1964b56b3106c0526ad2666e3dbcd9518597ccfc50d7e9de846b33e786398ff3

Observation a67a9fdf-6975-4f4b-8497-98a52fbdd8aa · inbound

Establishing Best Practices for Building Rigorous Agentic Benchmarks cites this paper.

Establishing Best Practices for Building Rigorous Agentic Benchmarks CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:28.958503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:28.958503Z digest=sha256:686c56cdc17fffa36d4b29de56f25140233fddee9320bbd144f9d0da410ea24f

Observation 8fa835e2-666e-4b34-b627-ebd4ee5cfe18 · inbound

FaultLine: Automated Proof-of-Vulnerability Generation Using LLM Agents cites this paper.

FaultLine: Automated Proof-of-Vulnerability Generation Using LLM Agents CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T15:43:01.005369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:43:01.005369Z digest=sha256:8ad37fd8ab20fcbc70b5d506f9e596c1286d4d2457dcba2314cfa80fc90714dc

Observation 3ef074a3-eee1-4112-841a-28577879b915 · inbound

Agent Identity Evals: Measuring Agentic Identity cites this paper.

Agent Identity Evals: Measuring Agentic Identity CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.658280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.658280Z digest=sha256:8bc2edf376c56af393851ed75e4656601fe0746125de7212eaa7d45ba0c4da56

Observation 79a15dea-2645-4b32-a23a-6c305b891379 · inbound

Running in CIRCLE? A Simple Benchmark for LLM Code Interpreter Security cites this paper.

Running in CIRCLE? A Simple Benchmark for LLM Code Interpreter Security CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T14:25:15.865587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:25:15.865587Z digest=sha256:93701aff9b687e8779b57b118e32fd3c3b15ebb0630f643bd8c7eacd05865ab6

Observation beaa2fdd-ee90-4065-92c4-86d51353daac · inbound

Reliable Weak-to-Strong Monitoring of LLM Agents cites this paper.

Reliable Weak-to-Strong Monitoring of LLM Agents CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T15:53:56.525194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:53:56.525194Z digest=sha256:ae8085189aca715054abb671704a00a24df950a8e08e503c3eb42b0192935617

Observation 97f5c947-d833-41c0-9eea-6a0424322ccf · inbound

xOffense: An Autonomous Multi-Agent Framework for Penetration Testing with Domain-Adapted Large Language Models cites this paper.

xOffense: An Autonomous Multi-Agent Framework for Penetration Testing with Domain-Adapted Large Language Models CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:41:38.180046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T16:38:45.171505Z digest=sha256:31fea8db01c66415c468cc21a6c491f56cc6d6e86c5b08c5b47d9d1f4bcd794e

Observation dc0e303d-5e45-4ea3-8856-20470f1765c2 · inbound

AXE: Grey-Box Exploitability Confirmation for Localized Vulnerability Reports cites this paper.

AXE: Grey-Box Exploitability Confirmation for Localized Vulnerability Reports CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T23:18:01.261262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:18:01.261262Z digest=sha256:5538c0b7821739fe399d7536d05fbf30179e5a7b1ccbf3c55ac107aaabb7ceca

Observation 29017f42-d4c8-4293-8c6b-9861a70b95d4 · inbound

RealVuln: Benchmarking Rule-Based, General-Purpose LLM, and Security-Specialized Scanners on Real-World Code cites this paper.

RealVuln: Benchmarking Rule-Based, General-Purpose LLM, and Security-Specialized Scanners on Real-World Code CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:40:23.968189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T12:39:37.047074Z digest=sha256:b97f0f69e6beb1360b179ebd96449de3dceae36a0e04a377022bde625225a808

Observation c46f35e0-fd53-4d56-89b6-7eac907f5e8c · inbound

CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios cites this paper.

CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:10:57.379648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T01:54:46.391345Z digest=sha256:2a5d51837eb476f1ba47b24e62806922e7ba10e8a8a87ce1cf3d33ff00aa66ec

Observation dbd44608-ca5b-447a-a142-76404729d276 · inbound

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World cites this paper.

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 51

Resolution
malformed identifier
arxiv_id, observed 2026-05-12T07:16:27.897902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T03:34:55.538935Z digest=sha256:7240f7e5bc1ca2b9f9872c8f3bae6e5b66cb60241641845b987d448dfb1d7a84

Observation 6e704e20-3f73-465b-9d9d-bed00b522524 · inbound

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World cites this paper.

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 50

Resolution
malformed identifier
no resolver link, observed 2026-08-02T14:22:27.799890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:22:27.799890Z digest=sha256:d2d5b3426b5e010e329de8b7c746e693058a8ed2b2deade16a17b85135bfb37e

Observation 2f2d7780-46d2-4354-b85c-44302813bbee · inbound

CyberEvolver: Structured Self-Evolution for Cybersecurity Agents On the Fly cites this paper.

CyberEvolver: Structured Self-Evolution for Cybersecurity Agents On the Fly CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:23:58.967193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T21:19:59.005348Z digest=sha256:8b4580ea1374847c74c8821266bd12943b8b1b7ebf4f1607f2eaebe1367123f3

Observation 562dc6d0-c2ee-4135-8e6b-d1fa3c1bedb6 · inbound

FORGE: Multi-Agent Graduated Exploitation and Detection Engineering cites this paper.

FORGE: Multi-Agent Graduated Exploitation and Detection Engineering CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:46:32.336920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T09:45:06.217817Z digest=sha256:76fd4d499a05e4ed903613e066782f4db6e7299c8af65d7d1c2a44bb324df487

Observation c39a3bd8-9e5f-4367-819e-4106573ee9d1 · inbound

ZERO-APT: A Closed-Loop Adversarial Framework for LLM-Driven Automated Penetration Testing under Intelligent Defense cites this paper.

ZERO-APT: A Closed-Loop Adversarial Framework for LLM-Driven Automated Penetration Testing under Intelligent Defense CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:16:58.745781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T01:26:22.434043Z digest=sha256:52d4e6dc08d3925c40b9ffdb58eba3d2fde20575239bde504a296ee97a1017ef

Observation a0c73dcc-c466-4dc4-a6b4-84e6645bdcfe · inbound

The Emergence of Autonomous Penetration Capabilities in Large Language Model-Powered AI Systems cites this paper.

The Emergence of Autonomous Penetration Capabilities in Large Language Model-Powered AI Systems CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:38:33.539484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T06:25:17.298114Z digest=sha256:3aaee5e8db087814cdcb6f0da692c7fb5c5246c2ee9c2ae9f7546ddfdb1df7f2

Observation d686cf9b-36b1-4420-b43f-00d903c08b40 · inbound

The Emergence of Autonomous Penetration Capabilities in Large Language Model-Powered AI Systems cites this paper.

The Emergence of Autonomous Penetration Capabilities in Large Language Model-Powered AI Systems CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:34:38.128616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T11:27:22.436981Z digest=sha256:a82cde4814fce21c4ce08b0902be7afb27f8ea602a9f8fbf267686f40613e0d0

Observation 8299dd19-e561-47ad-b4a6-b5f4f8879a2a · inbound

AgentCyberRange: Benchmarking Frontier AI Systems in Realistic Cyber Ranges cites this paper.

AgentCyberRange: Benchmarking Frontier AI Systems in Realistic Cyber Ranges CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:48:40.388696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T04:58:34.360751Z digest=sha256:0e40197c31482d0cc5b579cd98d3e2d9f98a8354f445f500d6b9a06ce3df0623

Observation 80398012-910e-4a7d-ad39-d1e97d416618 · inbound

Certifying Ghosts: How Cybersecurity AI Agents Break the EU Cyber Resilience Act cites this paper.

Certifying Ghosts: How Cybersecurity AI Agents Break the EU Cyber Resilience Act CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-09T20:16:29.520137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T20:05:20.263249Z digest=sha256:5c249f574388746d98315d3b399f774929ca33598fe4713e253868f592563e2b

Observation 00a49258-7ad4-4e0c-8a36-163edbb7f068 · inbound

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents cites this paper.

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T23:44:37.460029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:44:37.460029Z digest=sha256:89a3b1770e009cfbad06efcdd4e5f558f40cc31134ae8c967689611718c0683c

Observation 27b9270f-cf53-42fd-96ab-6c1e058371f5 · inbound

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction cites this paper.

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T15:36:02.680411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:36:02.680411Z digest=sha256:bac0c115388f584a83057b704ad998eb16f30550ef32e26e74eb6d8de4b9055c

Observation 3091db84-ccdc-430a-9224-a1055915bee6 · inbound

Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response cites this paper.

Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T02:40:29.390089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:40:29.390089Z digest=sha256:27f51d04e6f09affabb4c619601f1468fa66fa96bfec2ed35f2c3402bb495d31

Observation 77fa3257-b6a1-4c4b-96ae-57bdf59e429c · inbound

Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response cites this paper.

Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T01:33:20.873642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:33:20.873642Z digest=sha256:fcbdb816f3b9ce05074b87ef68738a3cdde8b8e8bd875d9c351f0f1cd515abff

Observation da23d235-37b0-4ae6-87df-18657f435f7c · inbound

SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response cites this paper.

SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-30T21:00:41.529969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T21:00:41.529969Z digest=sha256:2f19f8ae73577313d32aba351c5760e9d2c53ec3e8dcf5971118fbf8d2272101

Observation 6064daf8-660e-4d84-9fa8-952968d17caa · inbound

AgentSnare: Learning to Delay, Divert, and Defuse Autonomous Penetration Agents cites this paper.

AgentSnare: Learning to Delay, Divert, and Defuse Autonomous Penetration Agents CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-30T14:30:12.522679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T14:30:12.522679Z digest=sha256:969130acc9baf0a14b1d0b3b7c307b11fd70d01ffb280ba18bc572ad294694ee

Observation 00732459-9b20-4e58-b20b-5bfeefefedea · inbound

AgentSnare: Learning to Delay, Divert, and Defuse Autonomous Penetration Agents cites this paper.

AgentSnare: Learning to Delay, Divert, and Defuse Autonomous Penetration Agents CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T04:29:02.210995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:29:02.210995Z digest=sha256:fa01b6beeed328fd44a0573ac3085f2ade46180448ac4c0661fc4745dee5e9eb

Observation 93a5464a-3e34-4fb2-80e2-97f2de3efcc9 · inbound

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts cites this paper.

From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:21.169998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:21.169998Z digest=sha256:dacd4ed5422305c6e59dc354c82290c0e188f8f5b0a0e4a52d1acc3930744fd0

Observation 868f2da6-28b3-4df1-bd39-6bb51e0ea382 · inbound

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark cites this paper.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T14:19:09.549877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:19:09.549877Z digest=sha256:9e0904d51ab552a7ec02b87177eb6804620c86a14a161e32d95e92dea957ce19

Observation 97d3c3fa-887f-4d8e-bf85-d9d2c3c98277 · inbound

Rethinking Agent Security as a Networking Problem cites this paper.

Rethinking Agent Security as a Networking Problem CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-16T00:16:52.653903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:16:52.653903Z digest=sha256:4173c67a954a95a3d560cc61fcb662c321cc631cfaab8b4948783032e0323631