Pith. sign in

Paper Citation Record · LEDGER

NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2406.05590.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.05590 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T23:10:11.091917Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 95d04523-eaaf-47c8-9261-b400cd3c580f · inbound

Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks cites this paper.

Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T23:10:11.091917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T23:10:11.091917Z digest=sha256:ee1c5fb5ac8549873157d1f12a963bb1fd99f15a0d97a3c592d041f829553b96

Observation d163595e-7ebc-4faf-a91d-e25a6264d473 · inbound

Improving LLM Agents with Reinforcement Learning on Cryptographic CTF Challenges cites this paper.

Improving LLM Agents with Reinforcement Learning on Cryptographic CTF Challenges NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:35.404055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:35.404055Z digest=sha256:d3041a37728ed291678ed66d94fb5e9ec73cbffc5bcc647110b94c8311eb5810

Observation 425b3577-9f55-4a6f-9635-ccaf8025aa32 · inbound

Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges cites this paper.

Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:05.819135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:35:05.819135Z digest=sha256:38ce29b0fb45c3efda75630c48e19375668ad0e117da44d48858b82fbc43ccb2

Observation 2c5271e7-fee5-4224-bf41-1af1cef21a82 · inbound

Running in CIRCLE? A Simple Benchmark for LLM Code Interpreter Security cites this paper.

Running in CIRCLE? A Simple Benchmark for LLM Code Interpreter Security NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T14:25:15.835767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:25:15.835767Z digest=sha256:c7f0ec38dd741152a4db6186736095df81d7661cbee362ca9039ae6bef3ca4a5

Observation 465b323a-806e-4e04-8059-d42cec2a3e14 · inbound

Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks cites this paper.

Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:06:19.098028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T06:02:32.399075Z digest=sha256:ae19436555b2e7ab0b3c01fce5398a22e13486cd76b15d812c5e73161cc8438d

Observation d0e9112a-cc10-45ad-9357-1898e1e3df85 · inbound

Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps cites this paper.

Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:06:02.871221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T02:23:56.777888Z digest=sha256:044951e81e3c5408130106ebef85a5c2f0d1d9205358464d070db980cc0dc0bc

Observation 495deee8-79ee-4084-8dc3-77be9f21d781 · inbound

Synthesizing Multi-Agent Harnesses for Vulnerability Discovery cites this paper.

Synthesizing Multi-Agent Harnesses for Vulnerability Discovery NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:19:46.433919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T00:17:26.218561Z digest=sha256:7ae2d5db69f3f8554dabf42855c682e5cc5aa90e10e7400df2076ceccf22d5b4

Observation 96ce393e-2f4d-4192-a244-09cc28ad4275 · inbound

Dynamic Cyber Ranges cites this paper.

Dynamic Cyber Ranges NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:16:52.890773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T03:04:03.611481Z digest=sha256:f23815e1d3cf77bf130bf1154f5e91f8db893ff26b06d6a4dfecfc03a2b262bd

Observation fd25a401-3a2c-4e9a-8080-d2b893e40381 · inbound

Are Frontier LLMs Ready for Cybersecurity? Evidence for Vertical Foundation Models from Dual-Mode Vulnerability Benchmarks cites this paper.

Are Frontier LLMs Ready for Cybersecurity? Evidence for Vertical Foundation Models from Dual-Mode Vulnerability Benchmarks NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:30:20.213891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T04:27:57.881555Z digest=sha256:4123d7f8ae02ca3931a959eea2604986b773673a4438debabc363efcee1bf45e

Observation 5d93b70a-bdf3-4a88-b9f4-25f4fbb1f84f · inbound

Are Frontier LLMs Ready for Cybersecurity? Evidence for Vertical Foundation Models from Dual-Mode Vulnerability Benchmarks cites this paper.

Are Frontier LLMs Ready for Cybersecurity? Evidence for Vertical Foundation Models from Dual-Mode Vulnerability Benchmarks NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:35:12.937455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T16:25:47.480189Z digest=sha256:6eb0b44d3dd69828d52d57c11be9e2b184b2cd92e1246f1fb60aff4575ac4b6a

Observation 916b8740-ad9a-4ba6-8d4b-655ba5d65ae2 · inbound

CyberEvolver: Structured Self-Evolution for Cybersecurity Agents On the Fly cites this paper.

CyberEvolver: Structured Self-Evolution for Cybersecurity Agents On the Fly NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:23:58.993128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T21:19:59.005348Z digest=sha256:0c7f6ef8873e0892bd8dbbdf86a239c94dbedaa391e2b597be6a465605b8b1c5

Observation 38612fa6-7e8f-427d-859d-f1d3a5b99f5f · inbound

unix-ctf: Procedural Environments for Unix-Competence Reinforcement Learning cites this paper.

unix-ctf: Procedural Environments for Unix-Competence Reinforcement Learning NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T11:23:20.894545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T11:19:38.959705Z digest=sha256:5e46640fdfb41c751f08172de5c8d925b596ef5adcaf7afd7afe28f90a5c2954

Observation 58921728-1b32-4b4e-a8a8-94aedc70557c · inbound

Mastermind: Strategy-grounded Learning for Repository-Scale Vulnerability Reproduction cites this paper.

Mastermind: Strategy-grounded Learning for Repository-Scale Vulnerability Reproduction NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:08:21.515900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T14:02:06.173081Z digest=sha256:ca22fb82345dd4705e4f35a8e0a8e3da0dc6a5d9967ed29bce97e46c885e3a97

Observation 8d20a8c4-390a-4650-a315-0bcba00456e7 · inbound

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents cites this paper.

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T23:44:36.919216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:44:36.919216Z digest=sha256:bfb3145642fba7e33b3fd02cc4aa442581ecb1e6bf2cfdcc34d88368dea26d69

Observation 56e77957-a8b6-4938-9fce-98f4ecdf0eee · inbound

The Disruptive Impact of Large Language Models on Capture the Flag Competitions and the Path Toward Fair Play cites this paper.

The Disruptive Impact of Large Language Models on Capture the Flag Competitions and the Path Toward Fair Play NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T02:31:25.876707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:31:25.876707Z digest=sha256:b6a03010050e993a6310665d3bc47c3c7abfd56f6cbb44206a329f3b191af1a5

Observation ff095c14-b3bc-4ac3-910d-20329497c5a8 · inbound

Open Security Benchmark: Towards Autonomous Enterprise Cyber Defense cites this paper.

Open Security Benchmark: Towards Autonomous Enterprise Cyber Defense NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T10:23:19.237266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:23:19.237266Z digest=sha256:096f98fa309684426b438912bd54643a24d3190f024acb44a4995ae22b0753c4

Observation b62f91e6-9de3-4e1c-8460-edc37fea8a3c · inbound

Antares: Foundation Models for Agentic Vulnerability Localization cites this paper.

Antares: Foundation Models for Agentic Vulnerability Localization NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T07:50:50.682477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:50:50.682477Z digest=sha256:71803a82ac19778aee9772a0da3a7bcc948d811a2f64cbf7e2256b24946f8707