Pith. sign in

Paper Citation Record · LEDGER

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing

As of 26 July 2026, this Paper Citation Record lists 100 of 138 outbound references and 1 inbound Pith citation observation for arXiv:2604.05719.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.05719 v1

Coverage vector

measured 100 of 138 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T18:52:57.225878Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-26T06:30:07.085553+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T09:10:11.585499Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 138 outbound references displayed

  • verified exact21
  • verified fuzzy76
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8b511081-5ba5-4848-a39b-b08ac41d67ef · outbound

This paper cites https://openstd.samr.gov.cn/bzgk/gb/newGbInfo?hcno=BAFB47E8874764186BD B7865E8344DAF.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing https://openstd.samr.gov.cn/bzgk/gb/newGbInfo?hcno=BAFB47E8874764186BD B7865E8344DAF

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.922338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:591c9e634567a5a672b42d22d75f74f5d4f4f932f0af215b21fc6f0a06ca3f65

Observation ced6f320-887e-49de-88a4-d2b7c49063ba · outbound

This paper cites HexStrike AI MCP Agents.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing HexStrike AI MCP Agents

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.938328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:35349b8456b7bbf8a9ef1069e0571f445b2e5eb8e52c234dc9333a151e7aa42f

Observation 08a19b75-ef73-4418-9b25-e93e37403ef3 · outbound

This paper cites EnIGMA: Interactive tools substantially assist LM agents in finding security vulnerabilities.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing EnIGMA: Interactive tools substantially assist LM agents in finding security vulnerabilities

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.909426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:d0437675d2bcdfce34d739844c3687f96c07d2d4472af9f2c3ed01ceb2f0e4c5

Observation e54c609c-f1cd-4db5-b1e1-22c4981be066 · outbound

This paper cites Metasploit penetration testing cookbook.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Metasploit penetration testing cookbook

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.917910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:ef57d957a11b6737cb9806d01fa7943aa127eb8eb5cabaaecf4604b03a0ef9e4

Observation 7ad02f1d-8813-47c1-b803-218f81685de5 · outbound

This paper cites BreachSeek: A Multi-Agent Automated Penetration Tester.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing BreachSeek: A Multi-Agent Automated Penetration Tester

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:52.827605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:9823c349ece88e7eb36dc0a70a01eab6401345b86d2de1599713e577a407ca6e

Observation 1a4a38eb-8d14-4cf3-9467-0e9a30320173 · outbound

This paper cites Introducing the model context protocol.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Introducing the model context protocol

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.924718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:70ef9bd1580086b2ae21725f2dc4ff813b00aa654cfa428edd58946cb131da9a

Observation e1eaed53-6bff-43da-8d0d-a69d00170f57 · outbound

This paper cites Agent skills.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Agent skills

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.913453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:b82ea68f704659a7996991933e4e27ec410f36e5cad6cb3c4567247e6f6a24a9

Observation 4681805f-5208-430b-814e-4b77b21deb7c · outbound

This paper cites Claude code.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Claude code

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.907004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:c435e241f94f2ea9dc8dedac8f66145e2dc30cff1fc3f8e360fa28def78fc0f4

Observation e2ce3f8d-c201-4d37-9e4a-8cd6f3af2364 · outbound

This paper cites Claude opus 4.6 system card.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Claude opus 4.6 system card

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.936155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:c564b0c34df56d1bbdde7a83f0e628ac3e62503c7a74380aa93d6edba146efd9

Observation 4829b7bb-a499-47b1-95b3-d4329b765ac4 · outbound

This paper cites Introducing Claude Opus 4.6.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Introducing Claude Opus 4.6

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.919830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:1d944f4a86a35c643fadca1e418bfc7cd2089294f850fc2d71539be086c3d956

Observation df9bb2c7-d8f5-4857-b04a-cf3c90dc6b4d · outbound

This paper cites Software penetration testing.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Software penetration testing

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.911391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:5f896560188f82d7c2cb4d01cb7313b4cf3c351b5599b8b4e8aa6fc0b1013f6d

Observation 9bfcb707-cc4c-42c6-9ec4-fa9ae02499ee · outbound

This paper cites Pentest-ai, an llm-powered multi-agents framework for penetration testing automation leveraging mitre attack.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Pentest-ai, an llm-powered multi-agents framework for penetration testing automation leveraging mitre attack

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.927615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:811187d4f02bcbb46a689d2214911d9231ec19bd5fc790680b759bec054e7f89

Observation 4fe81752-659b-4afa-a2ec-b09ee134a740 · outbound

This paper cites About penetration testing.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing About penetration testing

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.933689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:89039052e20b52cd8d446dfd6d3fe6d6b043219c68db439b669b12aabf1b0373

Observation 83c95459-01d8-42fd-87d1-08833cb7faba · outbound

This paper cites Coverage-based greybox fuzzing as markov chain.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Coverage-based greybox fuzzing as markov chain

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.915871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:f05f1686d36133ccb96bee6f472d43ce586afb7aa98a8bfd27543e9e721cba59

Observation 8b19fa8d-b63d-4098-80d7-651dfd3b9094 · outbound

This paper cites Language models are few-shot learners.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Language models are few-shot learners

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.930889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:0bead7c25329ac7c9ecb8b79f28488372897e0217100b1d99ca77d62b281d1dd

Observation 735af62d-a28b-4649-b203-b74839e71d3c · outbound

This paper cites The diamond model of intrusion analysis.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing The diamond model of intrusion analysis

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.798392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:d53b5088f6b4eeef9e3a079cdf59f228807f1dc48ada99252903a4733ad04067

Observation 25b37231-0e44-4535-95a4-1520382eaaf7 · outbound

This paper cites an unresolved cited work.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Unresolved cited work

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.796193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:89eaddc83586b8b8094c38d8469ce04a48f9c9155d116a92781afed120c4b17f

Observation b3f955ad-ecab-4226-bf03-73196752c8a2 · outbound

This paper cites Why Do Multi-Agent LLM Systems Fail?.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Why Do Multi-Agent LLM Systems Fail?

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:59.242121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:f5b524ee4876544e5ec15702c84880c9387bc90c29fb088f6caab8da83509137

Observation 08536ac8-5ba6-4e02-b6d5-655d96d0404f · outbound

This paper cites tinyctfer.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing tinyctfer

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.697301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:e13e6666e0ec818e51c216e98c9ae77268c478c94e334a8528190ab610984062

Observation 2ed5c236-14f4-491c-92de-16f4d088dbd8 · outbound

This paper cites RedTeamLLM: an Agentic AI framework for offensive security.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing RedTeamLLM: an Agentic AI framework for offensive security

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:52.800364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:56ef4a8984be60442ec78a7c3f41a074aac8d6775ec24706c5597dd1e2655ed2

Observation 05c53256-eccc-4286-b99a-46af9ad72b53 · outbound

This paper cites Under the hoodie: Lessons from a season of penetration testing.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Under the hoodie: Lessons from a season of penetration testing

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.949648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:3128ece6099132531fc6fa7a5cafa5c840a7849de4dab411b30461b30bda4482

Observation d1856f37-ba71-46fa-905a-ba4c8f9b76db · outbound

This paper cites The growing importance of exposure management: Key insights from gartner hype cycle for security operations 2024.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing The growing importance of exposure management: Key insights from gartner hype cycle for security operations 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.826185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:e586b60679b5b9eaab8b3c092384eedfee1de9c28335205bdb93c1d82453cdaf

Observation 8d8c929c-c461-4124-ba8f-f18b68a554fa · outbound

This paper cites crewAI: Fast and Flexible Multi-Agent Automation Framework.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing crewAI: Fast and Flexible Multi-Agent Automation Framework

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.738107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:8bc52be71e5762258bb0217775ca6ada0786bd5124a1b330af2889065b853423

Observation 15a98dff-b576-4f2a-a0fc-27c0a41c9b0a · outbound

This paper cites CVE: Common Vulnerabilities and Exposures.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing CVE: Common Vulnerabilities and Exposures

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.685544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:e158102cd08f610767b8f1e0758f3b770485a1fc50d664e9e042975e81fa9481

Observation afc8c333-0720-4114-b781-cef46f630004 · outbound

This paper cites RefPentester: A Knowledge-Informed Self-Reflective Penetration Testing Framework Based on Large Language Models.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing RefPentester: A Knowledge-Informed Self-Reflective Penetration Testing Framework Based on Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:52.952640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:fc37739341d50b5e1905b10259c0297835f55baae88a22a5554acc1d8d92f94a

Observation 74fb8b87-e886-448f-9553-ff1b42f0ea4b · outbound

This paper cites Multi-Agent Penetration Testing AI for the Web.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Multi-Agent Penetration Testing AI for the Web

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:52.773454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:284a8dc004b8a1340c27118fcb311b394edefb0cf6a5c43c192737d6a26a6457

Observation 477558b5-eb75-406b-a2f9-3e4d0b4a1788 · outbound

This paper cites What makes a good llm agent for real-world penetration testing?.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing What makes a good llm agent for real-world penetration testing?

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.735714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:ffbbc28f2648ca155b0192bb2abb8a3e4d48d722cd33cf99ff64f3deeb7ed1e1

Observation aa96cb8d-3d2e-49c9-a22f-3e5c82a909f5 · outbound

This paper cites {PentestGPT}: Evaluating and harnessing large language models for automated penetration testing.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing {PentestGPT}: Evaluating and harnessing large language models for automated penetration testing

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.972537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:1066f6846b8c5e7ef9249f1f54865fb1bfb77e6b54772cf6da1b0be879b8b33b

Observation e551424a-534e-4971-8508-8b1dc06f0240 · outbound

This paper cites Cyberstrikeai.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Cyberstrikeai

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.976586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:f1c803eff18f4f98e3cb6650eed649d3a8bf14dd4e7587ff6d28afdb859ea31f

Observation dd13d838-f950-491a-a37e-b9cbcda988bb · outbound

This paper cites Regulation (EU) 2022/2554 of the European Parliament and of the Council of 14 December 2022 on digital op- erational resilience for the financial sector (DORA).

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Regulation (EU) 2022/2554 of the European Parliament and of the Council of 14 December 2022 on digital op- erational resilience for the financial sector (DORA)

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.987071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:a04416c2806d2fad32100c9087bac83919d62691da3e0d728ab9e3e4a2db164b

Observation 73a5fa88-342c-4570-8797-d4143e7031c8 · outbound

This paper cites A survey on rag meeting llms: Towards retrieval-augmented large language models.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing A survey on rag meeting llms: Towards retrieval-augmented large language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.791451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:2f8b6ad1497807eb1ea4d659492af209ae961a94d7700a6d415763ec9e9f1a8e

Observation dfd86db7-4ebb-4620-bfac-b64a76345362 · outbound

This paper cites Llm agents can au- tonomously exploit one-day vulnerabilities.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Llm agents can au- tonomously exploit one-day vulnerabilities

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.768208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:66a108ffbbb83a215b0e3f9614c6f71d215dc02389227947e34ff7749f9c14fe

Observation c72fffc0-1a6a-4728-99d2-a6fbfe733fe8 · outbound

This paper cites LLM Agents can Autonomously Hack Websites.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing LLM Agents can Autonomously Hack Websites

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:45:52.756965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:11317a6c656c90e829ad632747c78bdd741f41220dcd8ecd90d27081e408c0fc

Observation 01e02962-7938-4873-94e6-5f261f9116f5 · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-10T23:45:52.764061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:835b9dd73a6d1956df8e9e21f567278c61cab65a61a5e580d4e7995b28c55f85

Observation aedab074-ba1f-4542-84ef-888942865865 · outbound

This paper cites PentestAgent.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing PentestAgent

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.802284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:19db92e0406820ee71444f1247557699dca039c77bc8551e3371bfbc427460d7

Observation 1f6bd849-b4b0-455a-a271-d1221823af45 · outbound

This paper cites Automated Planning: theory and practice.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Automated Planning: theory and practice

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.954806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:f9712cad36bb26280c696f939ad9f09a0a7cfb6d95e89caaa5b960fcfe89c101

Observation c080e9c1-1285-44e7-878a-27c9c2143f21 · outbound

This paper cites Autopenbench: A vulnerability testing benchmark for generative agents.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Autopenbench: A vulnerability testing benchmark for generative agents

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.844123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:06bfc81979fb3e71bb536389fc8115483350a8b52192cf9eec54d7116ffff2c3

Observation 1a128853-efde-4765-9d42-4daebe9fc288 · outbound

This paper cites Gemini 3.1 Pro.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Gemini 3.1 Pro

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.821382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:ec0bc70b828f14a2eeea58f66f152e08e1ba04a917770b2fea4524abda08b4de

Observation 3f683134-c97b-45d8-a750-4aea0f6df451 · outbound

This paper cites Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Deepseek-r1 incentivizes reasoning in llms through reinforcement learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.823878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:9966cfb17a42583f838db1cfb4657db1ac4b7817798f3e163b5688a670e801a7

Observation 82cea609-b96b-43d8-90c0-f3f9a415a8f2 · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-10T23:45:52.795747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:6a170c903dc7a441375a60209b7e2afc76e23a7df60e73a463ff27d310af0938

Observation 9ddd11c2-4c53-4ed2-bd86-41bbf13fb3e2 · outbound

This paper cites Hack The Box.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Hack The Box

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.773399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:e43c54b6206ac5b409ab3e7c7a2ca11304c7b3a5a8ae96c6bde0a79e98cff9f8

Observation 307fbf41-9b28-488e-aa08-ce594834e3ab · outbound

This paper cites Hacking articles.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Hacking articles

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.833284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:890a445795df4a6bf1c8d1f2fcce1b3154e7c2725d89bc3a479d4a3308f88857

Observation a95bcb4f-082c-48fb-bdbd-0e7706c7a9dd · outbound

This paper cites Getting pwn nd by ai: Penetration testing with large language models.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Getting pwn nd by ai: Penetration testing with large language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.730109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:5822c0ca8aecaad46c25cc57d9d602c64dcf222dacbc15a44a0a8563d763d870

Observation 92e71700-932e-425a-a430-4441a80ceeb9 · outbound

This paper cites Can llms hack enterprise networks? autonomous assumed breach penetration-testing active directory networks.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Can llms hack enterprise networks? autonomous assumed breach penetration-testing active directory networks

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.807046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:b5217fd7dd65b11063153075de8273ad0f34bedca1c43e8039c701527d7b62fd

Observation 5ff18c7b-91bd-477f-95d3-153df0404d8b · outbound

This paper cites On the Surprising Efficacy of LLMs for Penetration-Testing.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing On the Surprising Efficacy of LLMs for Penetration-Testing

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:52.913350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:41e8a048c83a789b2f957e295cfb69c17756a97bc92c369ad7f7923f02c24f5a

Observation 5b685132-b264-4697-aaca-25fb7aaf15d9 · outbound

This paper cites Got root? a linux priv-esc benchmark.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Got root? a linux priv-esc benchmark

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.707813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:4e242db4f4b75762d4d74a3de1536c1915b99f713aeb1a4e441990f1efaf66bf

Observation d5d0d05b-ae42-4395-a5f5-d7bb50cf264e · outbound

This paper cites Llms as hackers: Autonomous linux privilege escalation attacks.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Llms as hackers: Autonomous linux privilege escalation attacks

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:52.792002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:847ac8e4ddefe241a0c0c6e505380dd87819fe5e5efdc81ef89ee41681f64b85

Observation 25d6001a-2db5-4982-b399-c6a12613c1d4 · outbound

This paper cites AutoPentest: Enhancing Vulnerability Management With Autonomous LLM Agents.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing AutoPentest: Enhancing Vulnerability Management With Autonomous LLM Agents

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:52.782770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:044d05203d9b67acff310fd7af12363c5b342fee517b7a280a4ab34ac2c21051

Observation 1b17aab1-a4cb-4474-9fc9-922a03da6a85 · outbound

This paper cites H-pentest.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing H-pentest

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.719105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:1c6d9aa597635fa7bfdc8647589840036ecfdcae17ad357f25f2cc6991956d57

Observation 4f22fead-67db-4381-816f-5244dd31e525 · outbound

This paper cites Context rot: How increasing input tokens impacts llm performance.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Context rot: How increasing input tokens impacts llm performance

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.761579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:e5a1bd2104d7a2154811fad2adac607155c3588779860fe023fd84fb6ecf38b2

Observation c14b06d1-a65f-4136-865d-b0b58db96bdf · outbound

This paper cites Metagpt: Meta programming for a multi-agent collaborative framework.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Metagpt: Meta programming for a multi-agent collaborative framework

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.778255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:84ea253a35505b2049d13a8a3398c5530b33228270b223238b133ff31cda0628

Observation c1bb90ce-f5d7-4194-9440-4192aaf72ef6 · outbound

This paper cites Penheal: A two-stage llm framework for automated pentesting and optimal remediation.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Penheal: A two-stage llm framework for automated pentesting and optimal remediation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.754598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:891ddb9973f7705c097e89ff3882e501efd3260b87a9d7e32cc61fd704f5acc9

Observation 9132d79d-3323-4586-8ef8-e99c740d6cda · outbound

This paper cites Qwen2.5-Coder Technical Report.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Qwen2.5-Coder Technical Report

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-10T23:45:52.748765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:2be2601dbde127780a0918736ee3779de0af8c56c2a9140b1b26f2a681e608f6

Observation 7d3879fa-7c87-4ac0-a4f4-cafb7eced539 · outbound

This paper cites newmapta.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing newmapta

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.745400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:173ba37a10020ae922cbdf4ed126e633b19a3a29e53ca214d7002553673916df

Observation 7e2a4e60-cd4a-4499-a760-aba0b7f599d0 · outbound

This paper cites Intelligence-driven computer network defense informed by analysis of adversary campaigns and intrusion kill chains.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Intelligence-driven computer network defense informed by analysis of adversary campaigns and intrusion kill chains

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.964831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:71d74ca33c422f18cd974ebb27562294d0cc22c26bbb4da3350fe8863ad35f80

Observation 15635540-4f50-4d4b-925f-f73cdba19d3c · outbound

This paper cites Towards automated penetration testing: Introducing llm benchmark, analysis, and improvements.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Towards automated penetration testing: Introducing llm benchmark, analysis, and improvements

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.732953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:560f5b6f6985774b8e5a7f1617c5ba0d340f44be970a9edc314d0535a3c25fae

Observation 08b243d6-5102-4a75-983d-fb9500e9f5ba · outbound

This paper cites Measuring and augmenting large language models for solving capture-the-flag chal- lenges.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Measuring and augmenting large language models for solving capture-the-flag chal- lenges

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.961495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:b9da3d0c307f58dc6ca05b05f3733dad1535117884389b81c28ab6c7f93805b7

Observation 55e1c6d5-d05d-4cf4-97e7-4a77f4799ef2 · outbound

This paper cites Survey of hallucination in natural language generation.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Survey of hallucination in natural language generation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.966988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:2578926a7e6861f7868d955c11a41a5f0b5ee063de20dd2c469239314126c19d

Observation 313fe888-8bfe-48f9-8ec9-c9b48df9767b · outbound

This paper cites Sok: Agentic skills – beyond tool use in llm agents.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Sok: Agentic skills – beyond tool use in llm agents

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.978600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:68ce901d4fb5affb8adaaf3f0f49be94d32ea92f901f226efb1792723fde2d9d

Observation 90b6b10a-f210-45e5-9d39-4d0474bd60c3 · outbound

This paper cites SWE-bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing SWE-bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.695416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:4b96057ea9cb903a33a761a579318ff96bca80d7de84653777f16580f720b7b5

Observation 07ea5548-35e0-4710-a5d6-012c57fef10c · outbound

This paper cites Dense passage retrieval for open-domain question answering.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Dense passage retrieval for open-domain question answering

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.699805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:70d005790520d021223619c237d8140448d9ed36be2f262c78dd7da68463361e

Observation e5edc5e0-73cd-408d-b6e8-1282c5c8cd02 · outbound

This paper cites Metasploit: the penetration tester’s guide.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Metasploit: the penetration tester’s guide

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.782399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:176ded0a699d1988f29f1dbe66c93ed1dcf9786533e1749712b7a965de1947f2

Observation 337c991f-c664-4ff4-9f66-92a4f5d068fe · outbound

This paper cites arXiv preprint arXiv:2508.07382 , year=.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing arXiv preprint arXiv:2508.07382 , year=

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:45:52.744985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:00b51d8ab104a0fae36a2b413c73e5fd2c96926ac454ab280fb6cbac6fac04ac

Observation a600fa08-1553-40b0-99bb-001ea4d61aca · outbound

This paper cites VulnBot: Autonomous Penetration Testing for A Multi-Agent Collaborative Framework.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing VulnBot: Autonomous Penetration Testing for A Multi-Agent Collaborative Framework

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:52.753000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:b85f206461808db4a8e811aa8cdf05220a3100f429161f3d3d3fa040a69bfed0

Observation fd9323a4-041b-4fb8-aca6-5935541fc4be · outbound

This paper cites Se perspective on llms: Biases in code generation, code interpretability, and code security risks.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Se perspective on llms: Biases in code generation, code interpretability, and code security risks

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.693120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:68222cd5f1255c3a82ce4a36fc9966b6a3c3e8408ba85e22e8440b32db1bb09a

Observation 8578de78-3a28-44e2-9518-44dedc88a023 · outbound

This paper cites LangGraph: Low-level orchestration framework for building stateful agents.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing LangGraph: Low-level orchestration framework for building stateful agents

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.688368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:3623c61f7956a7399abef64942b37e12420f09a3f06b212ed2bb71b66249da45

Observation 560a5d76-53a6-4569-8903-84934ca8e720 · outbound

This paper cites Lost in the middle: How language models use long contexts.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Lost in the middle: How language models use long contexts

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.752509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:4d1069b00e2a506b9f279daa454e0734b9c3cb21486c469b3aedb20265f1dce8

Observation 808051ee-2187-47b0-8dc5-6209b888f61d · outbound

This paper cites Pacebench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities.ArXiv, abs/2510.11688, oct 2025.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Pacebench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities.ArXiv, abs/2510.11688, oct 2025

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:52.760810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:3d951f2ec4dc254eb92131e6a21d76e0ccfc50dad5f859f4220d9ae018fcd4ff

Observation 4c7ec2b4-fed1-4d34-897f-83a7f644b1a0 · outbound

This paper cites Yu, and Ming Zhang.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Yu, and Ming Zhang

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.816490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:eb504192de448f5be4e0a6c8ddfa4640e1049a625a7c199694de6ee85f1b7f48

Observation 24ee55ed-9801-40dd-9d49-21e2b1e09b42 · outbound

This paper cites xOffense: An Autonomous Multi-Agent Framework for Penetration Testing with Domain-Adapted Large Language Models.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing xOffense: An Autonomous Multi-Agent Framework for Penetration Testing with Domain-Adapted Large Language Models

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-05-10T23:45:52.866163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:706db38224d48dfbf8c35a0b916070b366b4445c19f36684e0aa5f4dfde6bdad

Observation 57a3d93d-51c1-4900-ae5c-a214e0bb8ce6 · outbound

This paper cites xbow-competition.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing xbow-competition

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.980347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:657f03049af8eed0bf1ccbdb8b43112e5053f531d1729a728f909594763e63a0

Observation 81de04e4-3411-4450-a2aa-2967e39c4ee8 · outbound

This paper cites Shell or Nothing: Real-World Benchmarks and Memory-Activated Agents for Automated Penetration Testing.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Shell or Nothing: Real-World Benchmarks and Memory-Activated Agents for Automated Penetration Testing

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:52.816819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:f42b418c69f4a7e8b4eb9b89bf76c23d2ba7d5944181ea813597059b8fa053a9

Observation 29fbf94f-0ec5-4fba-a4a5-c4cb4be3e5da · outbound

This paper cites Graphical user interfaces.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Graphical user interfaces

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.994363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:b7a1c3331bf632a8a40eebcf2bdefd39f3678050bd48e374b2453086a9a02c6e

Observation ae4ade74-0f14-461f-a995-cdf7a66b5d9b · outbound

This paper cites CAI: An Open, Bug Bounty-Ready Cybersecurity AI.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing CAI: An Open, Bug Bounty-Ready Cybersecurity AI

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:52.777811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:1d1c23fcb5e936e8f5adfa3a16f4e134807bdf142a658062718d18c1353c54ed

Observation f2586f11-f75b-468b-bedc-0b8860a1a230 · outbound

This paper cites an unresolved cited work.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Unresolved cited work

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.839497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:64b7ebea2b3dc6cddbb13e17507d674333eb0f865fa9da8bcffa58ef69ae3d3d

Observation 4f7150bb-3455-4d69-b156-1cf7f8d49865 · outbound

This paper cites CWE-Common Weakness Enumeration.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing CWE-Common Weakness Enumeration

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.800349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:55ebe97a275f4f5e38d8f323421c4a3e1a8b453ff813f52831008162f81d2524

Observation 2fa01074-e596-408d-ac40-5a8d047eb89f · outbound

This paper cites Kimi Code CLI.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Kimi Code CLI

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.784554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:0a87d1f36bf2582b9eca5c995a7e7042a046f97e7ca05c6b00f737e873b35cf8

Observation 11e04272-a99f-4219-9bce-73a6cd349f42 · outbound

This paper cites Penetration testing and ethical hacking services market size & share analysis - growth trends and forecast (2025 - 2030).

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Penetration testing and ethical hacking services market size & share analysis - growth trends and forecast (2025 - 2030)

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.992249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:16907302fdb9e53d52d5124989de6d6f218dcc460cd523b2172a43cdd8cdc934

Observation ac72b7d5-c36c-4076-9c24-d44a01e37ff5 · outbound

This paper cites HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:45:52.946097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:6e4113d1c0648a2c1c491687b8355d6ce4d12add8c1be4b6c1d17f29d370ec99

Observation 43bd97bf-c2b2-4b8b-b7aa-1cef2b5ea14e · outbound

This paper cites Rapidpen: Fully automated ip-to-shell penetration testing with llm-based agents.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Rapidpen: Fully automated ip-to-shell penetration testing with llm-based agents

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:52.889783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:36ecc4d9d5d0914c3ac1b16e684e6ae1d85ede386c9cff7e1ee344f63fd2b944

Observation 3437b781-ae95-4d11-8d11-4210f02e680c · outbound

This paper cites ARACNE: An LLM-Based Autonomous Shell Pentesting Agent.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing ARACNE: An LLM-Based Autonomous Shell Pentesting Agent

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:52.854441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:8aa1c907916ef708224cbad9f1563fad05cd551166241b128c4cb54795721979

Observation e80334b4-8736-4575-9b26-0192498cdf6b · outbound

This paper cites Passage Re-ranking with BERT.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Passage Re-ranking with BERT

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:38:57.456949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:c12604659f424738f28409b7c66607b86a19f6bdd2cb0c0b5105716aaa85b433

Observation 56d2614a-dcc8-407e-ab80-2f59b5e508ec · outbound

This paper cites Introducing GPT-5.2.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Introducing GPT-5.2

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.787016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:fd880184366cb13a37c96c37967803019dab5f7bd358e3c7646942988b612ef3

Observation 1524d535-a30b-447a-87d7-98b7d19fda3e · outbound

This paper cites GOAD (Game Of Active Directory).

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing GOAD (Game Of Active Directory)

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.957288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:fb47eb5d47248f6f15f833c91963cf198e6c9e68e712246eff672e8feca405a6

Observation c986c91f-130c-4387-a2a0-cfe3665b7409 · outbound

This paper cites OverTheWire: Wargames.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing OverTheWire: Wargames

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.709904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:a26b1380830e646fb6e5b9fb2ba92550c9e21e53e3fa01b4d3e8fdc4b4adac7a

Observation 3efd8319-81cb-43ad-a734-2f54b633d769 · outbound

This paper cites OW ASP Top Ten Web Application Security Risks.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing OW ASP Top Ten Web Application Security Risks

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.742928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:47445dccfac673cecb8dc1ca8c9a74811d153bb637c3ab90c2b6fbbe14e9a53a

Observation ba9215ea-bc59-48bc-8229-268b351e6e72 · outbound

This paper cites Generative agents: Interactive simulacra of human behavior.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Generative agents: Interactive simulacra of human behavior

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.702770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:946aff4cf3f76cbe1a63d7e6e7574a4258472bea4fdc764613726cdcf067e18d

Observation 0505bcc1-83c5-4228-a86b-9c338242f49a · outbound

This paper cites ctfsolver.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing ctfsolver

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.722597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:b9cbc8821365d141ba028eb5b417ca7dca6d00d99b8ee3abb2c245fc1ee6ef12

Observation 2f7d2d1d-0e8f-42f9-868f-edfa3db28ddd · outbound

This paper cites PCI Security Standards Council.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing PCI Security Standards Council

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.793706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:49305c3fd3f0438be3c31dfe56723da0bb52c7087c95f52e88a461540d328247

Observation 249cb4d2-f8cc-48ec-8b86-ecc9739939ce · outbound

This paper cites The penetration testing execution standard.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing The penetration testing execution standard

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.947122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:4415b80d5e1e4f070d67ebd469fb6a1b1f9f5529d6f1801bb341c6d4f78cbe1d

Observation be7e0ebc-081a-40d5-baa1-b2d94f33344c · outbound

This paper cites Chatdev: Communicative agents for software development.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Chatdev: Communicative agents for software development

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.819287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:9fda31ad4021ddc3c609a1d8917a8e080496ccdb3b5e5c9f30ca8a8a517f0c90

Observation 019c8e3d-e4ca-4352-a12e-3f939e6593a5 · outbound

This paper cites Hackworld: Evaluating computer- use agents on exploiting web application vulnerabilities.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Hackworld: Evaluating computer- use agents on exploiting web application vulnerabilities

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.982490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:b06e3e99cd74c6dd037ac5bbc1fa9d68d0939108fa973b5b0dc2eca85526a967

Observation 517fae7d-fe0c-448b-88eb-8562dc1316c5 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Code Llama: Open Foundation Models for Code

Reference 93

Resolution
verified exact
local_arxiv, observed 2026-05-10T23:45:52.921947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:020fc44b5dddea4c3bc93fcccb395f12a65ce6285878717382b28933f9ec9aac

Observation 289c2aae-390f-4b8c-8afc-ac6b2e6201ca · outbound

This paper cites Luan1aoagent.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Luan1aoagent

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.846624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:e461eaba8165e568fee4e30ad82c26fe7c5fe882812e8e8f1fb5606985af2fcb

Observation 075b3a53-3cb6-4ed9-903b-08d7d659b540 · outbound

This paper cites Techni- cal guide to information security testing and assessment.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Techni- cal guide to information security testing and assessment

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.812007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:aef3174f7443a1a4877b9cef543e336e1491d81ad7ec566eac5d5538d4ab3579

Observation 3005aae5-69d8-4203-82f4-f5c97b708553 · outbound

This paper cites Toolformer: Lan- guage models can teach themselves to use tools.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Toolformer: Lan- guage models can teach themselves to use tools

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.809835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:56f1b41c3c53343fc64a3e1eb5e980cd6ffaf9b33c8f0e1a84d567d906ad0110

Observation 63ea598c-b462-46c4-bd7a-a290cee2202a · outbound

This paper cites An Empirical Evaluation of LLMs for Solving Offensive Security Challenges.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:52.935601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:55706ddb7ae5f3ab697b43cae3a18e77ca85cc4d461a5afc038cfe9bc1cf2cc0

Observation 36921e58-8a6b-468f-b8ba-acbcc7d8984c · outbound

This paper cites Nyu ctf bench: A scalable open-source benchmark dataset for evalu- ating llms in offensive security.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Nyu ctf bench: A scalable open-source benchmark dataset for evalu- ating llms in offensive security

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.968879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:985609347498f52d21157881b07381fdadef377a7d0118cd4fca5621b10a2a1d

Observation edde6d80-bd73-46d5-9e26-f33de730e60f · outbound

This paper cites Pentestagent: Incorporating llm agents to automated penetration testing.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Pentestagent: Incorporating llm agents to automated penetration testing

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.747708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:de4fef02114e135aef6599f50b8bcca59aaac6990c5c4bb5bebd2f8de8109a4d

Observation 839dae70-06ec-4174-9fdb-6537cdba8dfc · outbound

This paper cites Llms in software security: A survey of vulnerability detection techniques and insights.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Llms in software security: A survey of vulnerability detection techniques and insights

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.716041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:a153f4f5697d73daef5617c5f9c50bd45a3b47800dba17864849149e484e92cc

Pith citing papers

Observation d27b3552-3f0d-48be-862c-3c013b62a9e5 · inbound

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges cites this paper.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:64180474b1068cb76ff83bc9bd8ea4ea8f7fbaec8a2f0e805c76a5d38f19eaa5