Pith. sign in

Paper Citation Record · LEDGER

Autonomous LLM Agents & CTFs: A Second Look

As of 19 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 2 inbound Pith citation observations for arXiv:2605.21497.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.21497 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-22T01:06:49.411208Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T01:33:20.956536Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact1
  • verified fuzzy31
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 65631bbe-0d07-4a1c-a5bf-e7f12e1f1cca · outbound

This paper cites About penetration testing.

Autonomous LLM Agents & CTFs: A Second Look About penetration testing

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:10:52.621729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T01:06:49.411208Z digest=sha256:327548ee33803876553b8aeeff60db40b7dfb3151df95282121b7f41437ff67d

Observation 2a29ef05-ee5c-47a3-8e9b-b21577de53e8 · outbound

This paper cites Technical Guide to Information Security Testing and Assessment.

Autonomous LLM Agents & CTFs: A Second Look Technical Guide to Information Security Testing and Assessment

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:10:52.571147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T01:06:49.411208Z digest=sha256:9e2a0903b31c12a303d2da782a04d836057961a3ec5d2930fab1f6238533f636

Observation ca277ad3-b4c1-4e11-b5b4-bb702db98284 · outbound

This paper cites 2024 isc2 cybersecurity workforce study.

Autonomous LLM Agents & CTFs: A Second Look 2024 isc2 cybersecurity workforce study

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:10:52.589224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T01:06:49.411208Z digest=sha256:007f1f048bb07fcfe0abb0e13cefbc9975c27683e48ea9a10c8acd8525d9896c

Observation 19126d97-bfcd-4c41-a59a-54547ff168aa · outbound

This paper cites 2025 unit 42 global incident response report.

Autonomous LLM Agents & CTFs: A Second Look 2025 unit 42 global incident response report

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:10:52.618603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T01:06:49.411208Z digest=sha256:9d26ff55cdfe49c536468cb69c2a4c097580c238236cbd28dfd92dc94d29c822

Observation 5d16cdb7-eeb3-4d2d-835d-108d0d486319 · outbound

This paper cites When llms meet cybersecu- rity: A systematic literature review.

Autonomous LLM Agents & CTFs: A Second Look When llms meet cybersecu- rity: A systematic literature review

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:10:52.583544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T01:06:49.411208Z digest=sha256:23143d1c3ac7ba1c5999d2841a3e412923fcdaff4fbd71f992a04684f273de07

Observation afddec8c-61df-4bc0-b8b7-dc8c7c439c65 · outbound

This paper cites PentestGPT: Evaluating and harnessing large language models for automated pene- tration testing.

Autonomous LLM Agents & CTFs: A Second Look PentestGPT: Evaluating and harnessing large language models for automated pene- tration testing

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:10:52.625148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T01:06:49.411208Z digest=sha256:39c8255f8ce9b5796c64a34143302e3d13227f13ab897424d7a7c4eb114c26e1

Observation f6670d73-bf63-415a-a80c-89339dc98896 · outbound

This paper cites Getting pwn’d by ai: Penetration testing with large language models.

Autonomous LLM Agents & CTFs: A Second Look Getting pwn’d by ai: Penetration testing with large language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:10:52.574513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T01:06:49.411208Z digest=sha256:2518f9bb7f0851113055f3f91805bdc222bc394b00f4989b5a202f94ab954514

Observation 378fc359-2dac-4bd7-82a6-26a4d36fc448 · outbound

This paper cites Au- topenbench: A vulnerability testing benchmark for generative agents.

Autonomous LLM Agents & CTFs: A Second Look Au- topenbench: A vulnerability testing benchmark for generative agents

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:10:52.592722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T01:06:49.411208Z digest=sha256:5bb317897accf357a3d88204250fd5e5a048f4c26fb5f426dce917a2dfc2e52b

Observation b7b7df56-f1f5-4ba3-a0e1-9ae5c5050586 · outbound

This paper cites Teams of llm agents can exploit zero-day vulnerabilities.

Autonomous LLM Agents & CTFs: A Second Look Teams of llm agents can exploit zero-day vulnerabilities

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:10:52.580632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T01:06:49.411208Z digest=sha256:599f475a089047c3770c747b5c9259b2b10366c86c56d9c3c08758c80261024b

Observation 89369f9f-0ea1-480e-a635-3265117dd1fc · outbound

This paper cites Vulnbot: Autonomous penetration testing for a multi-agent collaborative framework.

Autonomous LLM Agents & CTFs: A Second Look Vulnbot: Autonomous penetration testing for a multi-agent collaborative framework

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:10:52.586421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T01:06:49.411208Z digest=sha256:3b52292c9667a9788c4b96bfb1b11b09676947716fbd2e4a0f001547a881593e

Observation cf585d34-426b-4744-90cb-8aa5135849b3 · outbound

This paper cites Multi-agent penetration testing ai for the web.

Autonomous LLM Agents & CTFs: A Second Look Multi-agent penetration testing ai for the web

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:10:52.577464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T01:06:49.411208Z digest=sha256:7f6e582629222dd99380f1870e3f420d94f364b7ce5903250fe66fcde122ce77

Observation 1e4a36e5-7367-4c5a-a9f0-03547a78bfe6 · outbound

This paper cites Cybench: A framework for evaluating cybersecurity capabilities and risks of language models.

Autonomous LLM Agents & CTFs: A Second Look Cybench: A framework for evaluating cybersecurity capabilities and risks of language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:10:52.613139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T01:06:49.411208Z digest=sha256:d8d150189778afedd23c44874add4f1f0debf81391a842380be2feec9b635159

Observation 77dd9b82-5446-4866-9498-6f403a3a61ff · outbound

This paper cites CVE- bench: A benchmark for AI agents’ ability to exploit real-world web application vulnerabilities.

Autonomous LLM Agents & CTFs: A Second Look CVE- bench: A benchmark for AI agents’ ability to exploit real-world web application vulnerabilities

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:10:52.543806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T01:06:49.411208Z digest=sha256:69dbc3f567b7de37abb079f767251d8d301ac1d9c6fc4df6faeef80081591b0d

Observation f73f074d-d939-4da0-a9a5-fa52cc40daa2 · outbound

This paper cites Claude is competitive with humans in (some) cyber com- petitions.

Autonomous LLM Agents & CTFs: A Second Look Claude is competitive with humans in (some) cyber com- petitions

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:10:52.555366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T01:06:49.411208Z digest=sha256:151afea27c42f35e8f63e66647eeb031e06c959322d706eceb54cd18bb7f2ea4

Observation ea6651d5-949a-42c1-aa92-684441332cd2 · outbound

This paper cites The road to top 1: How xbow did it.

Autonomous LLM Agents & CTFs: A Second Look The road to top 1: How xbow did it

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:10:52.540441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T01:06:49.411208Z digest=sha256:7cab97a25d607f5b7c1ca61496da350ee09612f80c6a27c750bb0ddebfe729df

Observation 0564a59e-536d-4e6f-bc46-61f526e91308 · outbound

This paper cites Comparing AI agents to cybersecurity professionals in real-world penetration testing.

Autonomous LLM Agents & CTFs: A Second Look Comparing AI agents to cybersecurity professionals in real-world penetration testing

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:10:52.550047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T01:06:49.411208Z digest=sha256:0a1d96db895d1b0ff2f9e4f51a99f171da03778fa45f75d082047f7dd4ca7646

Observation 25e5a18b-c12f-415e-bac0-b680178de25a · outbound

This paper cites Ten years of{iCTF}: The good, the bad, and the ugly.

Autonomous LLM Agents & CTFs: A Second Look Ten years of{iCTF}: The good, the bad, and the ugly

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:10:52.564671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T01:06:49.411208Z digest=sha256:0b9ef0c9f248bff6c58030df0435ffe138250d80980137c5e63f14c1fe3f65d3

Observation ec8ed76e-6561-4bf3-b3fc-7972d2f3019b · outbound

This paper cites Nyu ctf bench: A scalable open-source benchmark dataset for evaluating llms in offensive security.

Autonomous LLM Agents & CTFs: A Second Look Nyu ctf bench: A scalable open-source benchmark dataset for evaluating llms in offensive security

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:10:52.567948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T01:06:49.411208Z digest=sha256:a2671ba03f071e807bb7fb2aba2b87ada69b6edcb8467457c5f6c52d99e95beb

Observation ff9c4b1b-d053-40e5-86df-66464cc545c0 · outbound

This paper cites Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models.

Autonomous LLM Agents & CTFs: A Second Look Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:10:52.537401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T01:06:49.411208Z digest=sha256:6aa04d4abb20ad50398e7d3f5ae4eb1044078b7becf56092f896dcb96e1d38b3

Observation f01276ba-4259-4163-b670-d979c297c629 · outbound

This paper cites Evaluation and benchmarking of llm agents: A survey.

Autonomous LLM Agents & CTFs: A Second Look Evaluation and benchmarking of llm agents: A survey

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:10:52.534137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T01:06:49.411208Z digest=sha256:8e0c404884f49d0fb5ea0364a13f023935ae1792f8b02d40c5ab14f6f124c4fe

Observation e966506d-e06f-4fdf-9133-6ab1ae3f6719 · outbound

This paper cites Cognitive architectures for language agents.

Autonomous LLM Agents & CTFs: A Second Look Cognitive architectures for language agents

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:10:52.518454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T01:06:49.411208Z digest=sha256:81264a992d603f8378f73ce440138808de02688a0ff44ad173d808b4dfd133a5

Observation 2c2a69be-c7a6-4e40-bd09-980a684b4350 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

Autonomous LLM Agents & CTFs: A Second Look React: Synergizing reasoning and acting in language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:10:52.511806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T01:06:49.411208Z digest=sha256:26763483105e4f2e981023d54c98a3ff1e05b7cddcce0b444484d75bb2acc045

Observation 68c77307-cc5a-4205-94f4-8545165ba193 · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.

Autonomous LLM Agents & CTFs: A Second Look Toolformer: Language models can teach themselves to use tools

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:10:52.525097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T01:06:49.411208Z digest=sha256:710245421f0eb6001c6b45835d9ad4addca423bb198533e3e1baa626584673f5

Observation 7b065583-80c7-4932-b015-e43e5919f31f · outbound

This paper cites Autogen: Enabling next-gen llm applications via multi-agent conversations.

Autonomous LLM Agents & CTFs: A Second Look Autogen: Enabling next-gen llm applications via multi-agent conversations

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:10:52.521951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T01:06:49.411208Z digest=sha256:907b0748b62a12296c786ebd39a6cc3c8a2d42b2ad9ec0fc1c6cc9459487953c

Observation c42b2384-60df-470c-9250-727c08b33752 · outbound

This paper cites Agentverse: Facili- tating multi-agent collaboration and exploring emergent behaviors.

Autonomous LLM Agents & CTFs: A Second Look Agentverse: Facili- tating multi-agent collaboration and exploring emergent behaviors

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:10:52.561551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T01:06:49.411208Z digest=sha256:b014a74fd0cecb4b2a1d7af8c5cfc85b276f76865cc11a721db7aa9dd9d50dca

Observation 1f4960cd-ebbd-4799-8c85-423d66afe7b1 · outbound

This paper cites Cybersleuth: Autonomous blue-team llm agent for web attack forensics.

Autonomous LLM Agents & CTFs: A Second Look Cybersleuth: Autonomous blue-team llm agent for web attack forensics

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:10:51.529590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T01:06:49.411208Z digest=sha256:68cda18b892cc70f2fc8b2349ce5a6c29812c63c6eacb92f470a240f2f7451bb

Observation 22112cb9-b22f-4432-9470-f4f595dc790f · outbound

This paper cites From generation to judgment: Opportunities and challenges of llm-as-a-judge.

Autonomous LLM Agents & CTFs: A Second Look From generation to judgment: Opportunities and challenges of llm-as-a-judge

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:10:52.530996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T01:06:49.411208Z digest=sha256:b5bcb7a3c07e1d8d58d6d4ac613591065742ee707f07f185cf646d267aed9fc7

Observation 91567296-0566-40e7-8501-126d0b0faabf · outbound

This paper cites Claude code overview.

Autonomous LLM Agents & CTFs: A Second Look Claude code overview

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:10:52.528069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T01:06:49.411208Z digest=sha256:0a62abca2288f414b6681c19ec9d2455cc082231092e743fdd266de712000ccf

Observation b276f344-aaf2-462f-adaf-5eae590d3dc1 · outbound

This paper cites How claude remembers your project.

Autonomous LLM Agents & CTFs: A Second Look How claude remembers your project

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:10:52.515163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T01:06:49.411208Z digest=sha256:01c4661c10c0a2350b5cbe82c034449ae636f07f0e27cd4f3d6c78f15b6c8027

Observation f69ca69a-e98c-4edd-ba97-eec1c6c76ee0 · outbound

This paper cites Introducing gpt-4.1 in the api.

Autonomous LLM Agents & CTFs: A Second Look Introducing gpt-4.1 in the api

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:10:52.552575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T01:06:49.411208Z digest=sha256:6fd819347ddc9bcd52c5ae907791f2820f78eb1888b5e4e8acb7c1b4ec5c2da6

Observation fbc0ae28-1377-4e7b-a3b7-c63a0b2486db · outbound

This paper cites Gpt-5 system card.

Autonomous LLM Agents & CTFs: A Second Look Gpt-5 system card

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:10:52.546762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T01:06:49.411208Z digest=sha256:7f6890b14acaad7325c69d86db829b7612fab424afc7f32c182ca7f5ae440dbe

Observation cc2975c2-1c28-4ee3-9afb-70d514b3846e · outbound

This paper cites Claude opus 4.6 system card.

Autonomous LLM Agents & CTFs: A Second Look Claude opus 4.6 system card

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:10:52.558486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T01:06:49.411208Z digest=sha256:0d1ef49ca07ef3eb8023d94c401cce35eb2732c3620f789d4dcc964265789dcf

Pith citing papers

Observation fac8a509-d785-4d60-9a82-6f4937c6c2f0 · inbound

Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response cites this paper.

Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response Autonomous LLM Agents & CTFs: A Second Look

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T02:40:29.394859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:40:29.394859Z digest=sha256:9741d7d9e07a8910f851af396462d604ade9eeab5202443c47433f197f2b4697

Observation ac68106d-a200-4299-b430-4ecd33b67f01 · inbound

Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response cites this paper.

Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response Autonomous LLM Agents & CTFs: A Second Look

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T01:33:20.956536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:33:20.956536Z digest=sha256:6b57e80b7920da76c0f0607b99364bebb1f0c010ea8f4b47f4bf18a4bdaf047f