Pith. sign in

Paper Citation Record · LEDGER

An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2402.11814.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.11814 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:09:32.974241Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

6
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 50f8e8cb-e477-46cc-b372-20167b0f48ea · inbound

Psychometric-Based Evaluation for Theorem Proving with Large Language Models cites this paper.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.632242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.632242Z digest=sha256:eb1c1d0f16166ecc02c8cc33fd9c59247af9ff49a2eaa3d6f3359de8d1ae2fa6

Observation 666b8d88-e5e3-490d-a843-7faf46e44eff · inbound

Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks cites this paper.

Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T23:10:11.088564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T23:10:11.088564Z digest=sha256:7afcb01b2b1bec2303792dee70c914f732c8aa4692e20cdb7ce7d5f195e0475b

Observation 1650bf8f-41ef-4d5d-affe-6a3197337be7 · inbound

Large Language Models and Arabic Content: A Review cites this paper.

Large Language Models and Arabic Content: A Review An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T22:09:32.974241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:09:32.974241Z digest=sha256:5ce329f28ae1f2f9df911843daa6c905e2e13dc5137d1d277070376d42ed031a

Observation 433ca63f-fbdd-468a-886f-ec4a2eee13ca · inbound

AutoPentest: Enhancing Vulnerability Management With Autonomous LLM Agents cites this paper.

AutoPentest: Enhancing Vulnerability Management With Autonomous LLM Agents An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:15:52.720455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:15:52.720455Z digest=sha256:01013c7ae62eb9d019e9d3a83771810db91d8b6b58d978b7dc1555d2b4482943

Observation 25a9b8b4-5841-4955-bfe1-0d4179e79a7a · inbound

LLMs unlock new paths to monetizing exploits cites this paper.

LLMs unlock new paths to monetizing exploits An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T20:58:10.450150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:58:10.450150Z digest=sha256:bdb66921df026a1753f1ce5c669ef3c83231a7016b4619f010f8d3565a5205be

Observation 62362588-4ee9-4dd3-87b6-468428c4147d · inbound

CRAKEN: Cybersecurity LLM Agent with Knowledge-Based Execution cites this paper.

CRAKEN: Cybersecurity LLM Agent with Knowledge-Based Execution An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:51.648562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:23:51.648562Z digest=sha256:dfe5d588e2c2e9bff3ede32c1986d9a740cf84f9df87af6b75f35d75026f8916

Observation a2bdf34d-1c59-44e8-ba29-1d6dd7929034 · inbound

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models cites this paper.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.627164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.627164Z digest=sha256:e851e89559ee027d8789f0e71cc72c6569c9a4d4d005893969e79919957efa91

Observation 531bb9b8-a3ff-4145-b1e8-b789c3bd6c80 · inbound

Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges cites this paper.

Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:05.704955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:35:05.704955Z digest=sha256:cc2de7f676166f030f4738fea5acaac161a335b5ed3526677430f928800b08bb

Observation a0e3fa75-76a2-41c3-a90e-714a8f0405ac · inbound

On the Surprising Efficacy of LLMs for Penetration-Testing cites this paper.

On the Surprising Efficacy of LLMs for Penetration-Testing An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T21:10:06.656665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:10:06.656665Z digest=sha256:d18b74d49cb630ffca6079c22e79447e52126ff69b75312e90c85b98de22ca7a

Observation 63ea598c-b462-46c4-bd7a-a290cee2202a · inbound

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing cites this paper.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:52.935601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:8126ce7b9ef9e01f8175005b5449c7b4e34c0d70fea6bbd6ecc61e7e6ea3cebd

Observation e6a33eb4-e0dc-45c9-b173-a5f249f0ccd2 · inbound

CritBench: A Framework for Evaluating Cybersecurity Capabilities of Large Language Models in IEC 61850 Digital Substation Environments cites this paper.

CritBench: A Framework for Evaluating Cybersecurity Capabilities of Large Language Models in IEC 61850 Digital Substation Environments An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T19:35:44.724362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T19:32:57.161371Z digest=sha256:ef53d4e0026aeeaa8724744a56f8caf5afc1768ef5834cb4fd34814425256421

Observation 4bb081e1-779b-4042-a3d7-129c3a192a6f · inbound

Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks cites this paper.

Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:06:19.100715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T06:02:32.399075Z digest=sha256:939816adf5686de4b4c56b2ac0ee7ae11731e1ed96245b1bd8c915284cf74020

Observation 6de744e9-7328-4157-b677-b69d9a320542 · inbound

RAVEN: Retrieval-Augmented Vulnerability Exploration Network for Memory Corruption Analysis in User Code and Binary Programs cites this paper.

RAVEN: Retrieval-Augmented Vulnerability Exploration Network for Memory Corruption Analysis in User Code and Binary Programs An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:55:21.293998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T04:43:10.944954Z digest=sha256:7dcfac3892580c38f2ed4f914023c875c924ee5a59101fc40f5167d37491be55

Observation 5a02c696-be0b-43d4-851f-134fe6f7b267 · inbound

CyberCertBench: Evaluating LLMs in Cybersecurity Certification Knowledge cites this paper.

CyberCertBench: Evaluating LLMs in Cybersecurity Certification Knowledge An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:34:47.041453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T00:32:59.120345Z digest=sha256:e69f4bcfbc8670af88f6ab85d9b2550232cbcab793c37f5c5d11d2189fc412ca

Observation 11fde743-fb16-43d0-8e06-4a41b9ff1d5e · inbound

Benchmarking Large Language Models for IoC Recovery under Adversarial Code Obfuscation and Encryption cites this paper.

Benchmarking Large Language Models for IoC Recovery under Adversarial Code Obfuscation and Encryption An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:50:58.078458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T01:01:15.586948Z digest=sha256:776ea59a6b71aafa222e873617cf6f8d63d6f40e59d2782798a92a9c53b729cc

Observation 6f3341c2-a8fd-46a0-b129-cf9e57cf8280 · inbound

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges cites this paper.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:e42cfe4e68e093d50c9869e39aa609311805c53fb15bf02b943d4ce18149b581