Pith. sign in

Paper Citation Record · LEDGER

Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2504.10112.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.10112 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T23:10:10.965917Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:20:07.161491Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 36a33f6c-8109-4d23-bf1d-1af92c0df62d · inbound

Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks cites this paper.

Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T23:10:10.965917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T23:10:10.965917Z digest=sha256:4769f39325d305a636b728b712b6c71252eb14d82cbb59a6922a8df707e7d7d9

Observation 39cd706b-239e-4186-b0b0-bbd991284f8a · inbound

From Promise to Peril: Rethinking Cybersecurity Red and Blue Teaming in the Age of LLMs cites this paper.

From Promise to Peril: Rethinking Cybersecurity Red and Blue Teaming in the Age of LLMs Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:59.924739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:59.924739Z digest=sha256:2e820ed0873c7826791168d1924c117108b0304c24fceba0789882b246bdb4db

Observation bed81608-05d1-4c89-a81a-27894b52f624 · inbound

On the Surprising Efficacy of LLMs for Penetration-Testing cites this paper.

On the Surprising Efficacy of LLMs for Penetration-Testing Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T21:10:04.491227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:10:04.491227Z digest=sha256:6598928788a27b9319892d588de2bb5370f26785c9c08c21087bea14db73d830

Observation a7a0094c-8ea6-4482-9e2b-b5cff3cee040 · inbound

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World cites this paper.

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:16:27.628828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T03:34:55.538935Z digest=sha256:b1d4e597354aef4a96a5544f27ecbe08afe65e176cb6890035eef0e7b21f752b

Observation f27747f2-bda1-48df-a2ac-0167ef37d704 · inbound

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World cites this paper.

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T14:22:24.494937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:22:24.494937Z digest=sha256:43a069e52e62ae3251662e2dd8fff3b9237747f1daa2924c7ba8a10c555473a2

Observation a7931abb-6c44-44a5-b9f6-26956f8d76bf · inbound

How Reliable Are AI Attackers Against a Fixed Vulnerable Target? A 400-Run Empirical Study of LLM Penetration Testing Consistency cites this paper.

How Reliable Are AI Attackers Against a Fixed Vulnerable Target? A 400-Run Empirical Study of LLM Penetration Testing Consistency Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T14:13:30.540842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T06:53:48.841077Z digest=sha256:bacf2c9205751b4d154e4c87dd503a4319a8c30f06eec60457db08844f34b20f

Observation 965f875d-4656-4dcb-ba14-a9069b9c74c0 · inbound

ZERO-APT: A Closed-Loop Adversarial Framework for LLM-Driven Automated Penetration Testing under Intelligent Defense cites this paper.

ZERO-APT: A Closed-Loop Adversarial Framework for LLM-Driven Automated Penetration Testing under Intelligent Defense Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:16:58.787116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T01:26:22.434043Z digest=sha256:4b194a49ffe3cf4a02455a3802de8bfa1bccb81621597360e47539253dfb3e40

Observation cd011455-5dfd-41fa-a915-3f2ceec8281d · inbound

Decoupling Reconnaissance and Exploitation: Measuring the Capability Boundaries of LLM-Based Web Penetration Testing cites this paper.

Decoupling Reconnaissance and Exploitation: Measuring the Capability Boundaries of LLM-Based Web Penetration Testing Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:20:07.162950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T21:23:32.255067Z digest=sha256:12ffcccadce08b69adbdf3b2869f77b88f683afcb686e51afd6e51dd281f0802

Observation f4717228-edf6-44d4-814c-bec27062b365 · inbound

Hephaestus: Toward a Cybersecurity AI Scientist cites this paper.

Hephaestus: Toward a Cybersecurity AI Scientist Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T13:54:44.479468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T05:42:32.460183Z digest=sha256:ecb52d75e17dc611574e3fb1b87013fe0ec2774b4596d761d661b7c1df6113d6

Observation 5371b2c9-2f19-4d1c-b521-5fb7f4011829 · inbound

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges cites this paper.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:6db4c4da8f738aebbe6488327dba8e9989cb37d3a5a06cc04373615acb01d347

Observation 770e4217-b261-4293-ba14-82a3cefdce29 · inbound

Determinants and Limits of LLM Security-Tool Orchestration: A Study with HexStrike-AI cites this paper.

Determinants and Limits of LLM Security-Tool Orchestration: A Study with HexStrike-AI Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T06:28:52.906906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:28:52.906906Z digest=sha256:ba7dacb4d2feee8fb0d92b387811eaaebe601bc3fa4909a5034f8e6fb3469292