Pith. sign in

Paper Citation Record · LEDGER

LLM Cyber Evaluations Don't Capture Real-World Risk

As of 19 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 4 inbound Pith citation observations for arXiv:2502.00072.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.00072 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T22:04:33.593990Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T23:10:11.025662Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T02:07:33.281632Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact2
  • verified fuzzy22
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2b184fcb-075b-470a-a0cf-0d22dbc76a76 · outbound

This paper cites Introducing computer use, a new Claude 3.5 Sonnet , and Claude 3.5 Haiku , 10 2024.

LLM Cyber Evaluations Don't Capture Real-World Risk Introducing computer use, a new Claude 3.5 Sonnet , and Claude 3.5 Haiku , 10 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.663333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.328815Z digest=sha256:278854c7b57f9c35ec9c4b899135e5de67e877439a0266c24bf5a0623526b3ef

Observation fc76035d-be20-42f0-87af-4cb0ed18e822 · outbound

This paper cites Phishing Activity Trends Report , 4th Quarter 2023.

LLM Cyber Evaluations Don't Capture Real-World Risk Phishing Activity Trends Report , 4th Quarter 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.646352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.335713Z digest=sha256:a8cae4505a915f94ff346c22309afacdb2a7eb8937e596f2f1606d7d12040a1c

Observation 08c4b518-f91c-4c21-9ebd-b50c53109244 · outbound

This paper cites Phishing Activity Trends Report , 3rd Quarter 2024.

LLM Cyber Evaluations Don't Capture Real-World Risk Phishing Activity Trends Report , 3rd Quarter 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.628895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.341958Z digest=sha256:1fda0777395031045345c98a4fa4379130b71f29ec26c4e84b7efd0d8bdd5a9f

Observation fb3ca9dc-e754-4cb1-aafd-d21afdbda140 · outbound

This paper cites and Kogtenkov, A.

LLM Cyber Evaluations Don't Capture Real-World Risk and Kogtenkov, A

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.611472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.347979Z digest=sha256:35a629db03bbc7765b4c32d6c86cd06434b716f74012caece2cf75892f565931

Observation ed79ba16-725a-4a3d-8408-82246b218a83 · outbound

This paper cites Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities.

LLM Cyber Evaluations Don't Capture Real-World Risk Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.354128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.354128Z digest=sha256:f01425fdbbd74ef9a3c4e561d8eb3242347aa3ddd356113e66acd1fec32bb3fa

Observation 1d434117-3d9b-49e2-921a-e9a825aa9ec7 · outbound

This paper cites "Real Attackers Don't Compute Gradients": Bridging the Gap Between Adversarial ML Research and Practice.

LLM Cyber Evaluations Don't Capture Real-World Risk "Real Attackers Don't Compute Gradients": Bridging the Gap Between Adversarial ML Research and Practice

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.361159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.361159Z digest=sha256:93094111ee2be3eab64d37f35d32d0a27c720312ba19189d3480219c74dfa6ed

Observation a859d208-3314-48fe-8ea4-cb7d3c5f5364 · outbound

This paper cites CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models.

LLM Cyber Evaluations Don't Capture Real-World Risk CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.368368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.368368Z digest=sha256:300665e854fd2a887a866524669220bbcd96376b7b7c8f8ae002a5c6de429e5e

Observation 95c0f915-db67-47e4-ae77-5fb642fa5e28 · outbound

This paper cites From naptime to big sleep: Using large language models to catch vulnerabilities in real-world code, 11 2024.

LLM Cyber Evaluations Don't Capture Real-World Risk From naptime to big sleep: Using large language models to catch vulnerabilities in real-world code, 11 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.594239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.374456Z digest=sha256:9cf77bf0752e7f0dbe35506f0cc5a307be78ca37e845abefe862f1aaa5dd0287

Observation 8485747e-d971-41e3-9105-3c754be4f8ed · outbound

This paper cites The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation.

LLM Cyber Evaluations Don't Capture Real-World Risk The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.379294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.379294Z digest=sha256:de22694a24a4b5a558fbb60bc322627de3fa4209efe6c2364b5b1adc10fef16b

Observation 43d800fb-3ce9-4eb8-b9c0-5a955b284f57 · outbound

This paper cites FunkSec – alleged top ransomware group powered by ai, January 2025.

LLM Cyber Evaluations Don't Capture Real-World Risk FunkSec – alleged top ransomware group powered by ai, January 2025

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.577335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.385053Z digest=sha256:e4870eccd886373c55b9557f8c6ef9f47b799dffe26443b99e52355ee12601f9

Observation bc70c310-3b1d-48ab-9d0a-263ab40871b0 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

LLM Cyber Evaluations Don't Capture Real-World Risk Evaluating Large Language Models Trained on Code

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.390435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.390435Z digest=sha256:d75c23e49e4f3a8c2217d2c166387235de3cc59aa65fbbdb23a1ac00aa82f503

Observation 453d3976-525f-473d-8a7e-c3502375ead2 · outbound

This paper cites PentestGPT: An LLM-empowered Automatic Penetration Testing Tool.

LLM Cyber Evaluations Don't Capture Real-World Risk PentestGPT: An LLM-empowered Automatic Penetration Testing Tool

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.395712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.395712Z digest=sha256:7da78325c1ce6654f7e6909e1b433a57760aa4a0edbb0fafae3b1acfbb0bff89

Observation 210dca36-a026-428d-8629-726b72f7e63f · outbound

This paper cites Robust physical-world attacks on deep learning visual classification.

LLM Cyber Evaluations Don't Capture Real-World Risk Robust physical-world attacks on deep learning visual classification

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.560656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.401150Z digest=sha256:c997c313281c9df322c97afb25dfd44c4779c4ed8b96c84e42fa72ab10b7b3fe

Observation 63bcd5fb-1dcb-4b55-bca7-d0b4571c825c · outbound

This paper cites LLM Agents can Autonomously Exploit One-day Vulnerabilities.

LLM Cyber Evaluations Don't Capture Real-World Risk LLM Agents can Autonomously Exploit One-day Vulnerabilities

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.405915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.405915Z digest=sha256:e10e19ced5092209927e9fa30d3b01335447fe2a3c8e5cf78e2ac2e917704243

Observation c15c0154-c66e-497c-a7a7-88b441c96339 · outbound

This paper cites LLM Agents can Autonomously Hack Websites.

LLM Cyber Evaluations Don't Capture Real-World Risk LLM Agents can Autonomously Hack Websites

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.411892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.411892Z digest=sha256:bde6934ad2657d8aad5aeb8f8e8841cf645c32983795a75975f24cebf852c9e5

Observation fe24e7d6-45c6-4908-9b35-38e2300b9234 · outbound

This paper cites Teams of LLM Agents can Exploit Zero-Day Vulnerabilities.

LLM Cyber Evaluations Don't Capture Real-World Risk Teams of LLM Agents can Exploit Zero-Day Vulnerabilities

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.418057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.418057Z digest=sha256:1f3898d06f139fcf83699e9f8246a0628ce3322a302ce847cd6501a211d1ef3d

Observation cb522df9-8e68-4256-a28d-4503783d83f1 · outbound

This paper cites 2023 Internet Crime Report.

LLM Cyber Evaluations Don't Capture Real-World Risk 2023 Internet Crime Report

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.542751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.424270Z digest=sha256:afc7c0cefc7e75b1b8687980eb78fe1182667f06ca81bec0ade47ea64491465b

Observation e22701fb-69b4-4d82-b6fd-02173e688d60 · outbound

This paper cites BadLlama: cheaply removing safety fine-tuning from Llama 2-Chat 13B.

LLM Cyber Evaluations Don't Capture Real-World Risk BadLlama: cheaply removing safety fine-tuning from Llama 2-Chat 13B

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.429874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.429874Z digest=sha256:4bc930a7b965a3726233a20a02ff74f13c4612126d8ef6bf08ce1adc8dc5bee5

Observation 5c626d25-7818-471a-b928-31f4b8ba6b14 · outbound

This paper cites Sour grapes: stomping on a Cambodia -based ``pig butchering'' scam.

LLM Cyber Evaluations Don't Capture Real-World Risk Sour grapes: stomping on a Cambodia -based ``pig butchering'' scam

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.524676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.436838Z digest=sha256:d26ca91082c619d06acbe579422886f69fe40837334f07c917d3a2a92768c5c4

Observation 5e4213b4-87fb-45f4-8f6b-6a4dcc5b77cb · outbound

This paper cites Safety case template for frontier AI: A cyber inability argument.

LLM Cyber Evaluations Don't Capture Real-World Risk Safety case template for frontier AI: A cyber inability argument

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.442414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.442414Z digest=sha256:521618cb243844eb8d7844047aed30afd2ccab671fd7a82a609599618d3ec0ad

Observation b51324d8-4b85-41b8-900d-458036ad85ff · outbound

This paper cites Adversarial misuse of generative AI , 2025.

LLM Cyber Evaluations Don't Capture Real-World Risk Adversarial misuse of generative AI , 2025

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.505804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.448387Z digest=sha256:7fd4a1e04aa6be181da650b2f845f781146e0f030606000285237e5d462d0e43

Observation d423440c-c0c8-4d27-882c-bca395718d24 · outbound

This paper cites and Cito, J.

LLM Cyber Evaluations Don't Capture Real-World Risk and Cito, J

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.452963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.452963Z digest=sha256:38be06ad1460e1451a83814a4e8c3a14ec81815b5886dd06a21f8b56b984db92

Observation ba7e8388-6430-4110-b26f-cedd2ab27edb · outbound

This paper cites Spear Phishing With Large Language Models.

LLM Cyber Evaluations Don't Capture Real-World Risk Spear Phishing With Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.457578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.457578Z digest=sha256:55334fdd2dd316b7f053d5805e4fd9a47b4fd9226396645a3742752e05934249

Observation 5c2bca68-90e1-41da-8702-9d652720202e · outbound

This paper cites Evaluating Large Language Models' Capability to Launch Fully Automated Spear Phishing Campaigns: Validated on Human Subjects.

LLM Cyber Evaluations Don't Capture Real-World Risk Evaluating Large Language Models' Capability to Launch Fully Automated Spear Phishing Campaigns: Validated on Human Subjects

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.462290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.462290Z digest=sha256:de7fab7cdc1a2e45d11f373360f6094a99983237ea0f55af4bfe3db0c353106b

Observation d5bdeb83-b35b-4899-90dd-6e33de2dbc6d · outbound

This paper cites Devising and detecting phishing emails using large language models.

LLM Cyber Evaluations Don't Capture Real-World Risk Devising and detecting phishing emails using large language models

Reference 25

Resolution
metadata mismatch
raw_fallback, observed 2026-08-09T22:04:48.912079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.467563Z digest=sha256:9d8edb313479b91f00aeaa1c71766e368ebe9fd58fd61b60d384e2b7430e36ca

Observation db3a8d13-1121-4919-8dfb-e6678f9de43d · outbound

This paper cites An Overview of Catastrophic AI Risks.

LLM Cyber Evaluations Don't Capture Real-World Risk An Overview of Catastrophic AI Risks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.472242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.472242Z digest=sha256:f32926f895c008f0e238ebb565afc41ee67a0afcf3014c8a526af511b6581c8e

Observation 7c6542e2-94c5-4c4d-9565-6125617200a6 · outbound

This paper cites Why do nigerian scammers say they are from nigeria? Proceedings of the Workshop on the Economics of Information Security, 01 2012.

LLM Cyber Evaluations Don't Capture Real-World Risk Why do nigerian scammers say they are from nigeria? Proceedings of the Workshop on the Economics of Information Security, 01 2012

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.486659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.477435Z digest=sha256:eb4a74ddc8cfaa39a4ab67be6d511e69e89eeb508ea4f3b10debf77d781dd50c

Observation 0352e941-d60b-4224-b8ea-e544af213495 · outbound

This paper cites Now you see me, now you don't: Using LLMs to obfuscate malicious JavaScript.

LLM Cyber Evaluations Don't Capture Real-World Risk Now you see me, now you don't: Using LLMs to obfuscate malicious JavaScript

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.466991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.483374Z digest=sha256:bc5a1b21d15601e31867d557324348493a911d9bbec6a64193c648b6084c8a53

Observation 204d8de5-1490-4f04-a694-d6614981b0e3 · outbound

This paper cites Stored cross-site scripting (xss) in 2FAuth : How XBOW found a stored XSS in 2FAuth ( CVE -2024-52597), 12 2024.

LLM Cyber Evaluations Don't Capture Real-World Risk Stored cross-site scripting (xss) in 2FAuth : How XBOW found a stored XSS in 2FAuth ( CVE -2024-52597), 12 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.447646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.488615Z digest=sha256:5520a6b18144af47f62b192fb85d71ddbb744b0c10a672a8c014120419885ad0

Observation 27b08ecc-d8b3-4fe2-960e-f12d2f66672e · outbound

This paper cites Translated: Talos ' insights from the recently leaked Conti ransomware playbook, September 2021.

LLM Cyber Evaluations Don't Capture Real-World Risk Translated: Talos ' insights from the recently leaked Conti ransomware playbook, September 2021

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.424914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.494029Z digest=sha256:44005e8e58d5920fc096c1b355fd1a2effc206e8dbe106dea19b20a11f3ef3fe

Observation a4a79575-6512-4eeb-a6e1-73bcf464c3a5 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

LLM Cyber Evaluations Don't Capture Real-World Risk HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.499853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.499853Z digest=sha256:492cee297b431c072ad2f331d3c43f193c41b85db8b4beb805214ad600f1e337

Observation 12c1686f-abf5-498f-93b0-62c2c9ca7a6a · outbound

This paper cites An update on our general capability evaluations.

LLM Cyber Evaluations Don't Capture Real-World Risk An update on our general capability evaluations

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.405115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.505725Z digest=sha256:04888cb4423f8a20330c81dbb0e090e9b3897a748d51f3a5269279158392a88b

Observation ec9198f9-877b-451e-b4d4-c1d99862bdcd · outbound

This paper cites The Threat of Offensive AI to Organizations.

LLM Cyber Evaluations Don't Capture Real-World Risk The Threat of Offensive AI to Organizations

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-09T22:04:33.793851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.511132Z digest=sha256:12d4b924bd4201693822abe5f120b6482271930e05dd08c89716ac9045e03d9f

Observation c34d9cf7-baea-4763-8bc3-6a0ec7a7c574 · outbound

This paper cites MITRE ATT&CK.

LLM Cyber Evaluations Don't Capture Real-World Risk MITRE ATT&CK

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.385448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.516850Z digest=sha256:0354db025cdf2db2b77503b64e2e27eea41cd5bebf3015ebeb45066425a534a1

Observation 8b6af47e-590e-4d83-aa3a-ef7020aaee5c · outbound

This paper cites and Flossman, M.

LLM Cyber Evaluations Don't Capture Real-World Risk and Flossman, M

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.364755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.522871Z digest=sha256:01af13adf404fe64f0fc4aaf28ab17367b13794ab56151d18c445cdd03359768

Observation 42633144-4e19-4a8d-a742-cd32bce9bb40 · outbound

This paper cites Introducing Operator , 1 2024 a.

LLM Cyber Evaluations Don't Capture Real-World Risk Introducing Operator , 1 2024 a

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.344228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.528353Z digest=sha256:5eed3039cb117156926f33edd039796a80f5666cdd91e21d2ee300f438129bec

Observation 584f2ce6-6679-4b7b-a517-7f9fd7efe21e · outbound

This paper cites Disrupting malicious uses of AI by state-affiliated threat actors, February 2024 b.

LLM Cyber Evaluations Don't Capture Real-World Risk Disrupting malicious uses of AI by state-affiliated threat actors, February 2024 b

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.325794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.533729Z digest=sha256:0a20770a6038db5cf6153b66859188d818685ca3ecaea39cefb2052de53d2bcc

Observation d37c8ff2-dc4e-4943-acbb-6e223e991712 · outbound

This paper cites and Haugen, S.

LLM Cyber Evaluations Don't Capture Real-World Risk and Haugen, S

Reference 38

Resolution
verified exact
doi, observed 2026-08-09T22:04:33.640065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.539035Z digest=sha256:ddbb690110034a672b1b4129146ef24a00b0b36704cc2d6cf0f3449d66e67c96

Observation efc7f0e1-9d42-424e-89e6-55b39d2c85f9 · outbound

This paper cites No, LLM agents cannot autonomously "hack" websites, feb 2024.

LLM Cyber Evaluations Don't Capture Real-World Risk No, LLM agents cannot autonomously "hack" websites, feb 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.306353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.543901Z digest=sha256:fedbea94887da3572ef0d686c47efbcfb2800b455ebdfb04c32047bbab4a7b60

Observation 22c56035-901b-4691-9c58-eb886fe33a14 · outbound

This paper cites SoK: On the Offensive Potential of AI.

LLM Cyber Evaluations Don't Capture Real-World Risk SoK: On the Offensive Potential of AI

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-08-09T22:04:33.767100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.548908Z digest=sha256:5d85fb58b27dd01fd15d1df092654812eb305b5d676d8fe8fbbdb66addc41b52

Observation 04ff266c-6a39-49da-ba99-b6aab206b6b1 · outbound

This paper cites NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security.

LLM Cyber Evaluations Don't Capture Real-World Risk NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.554414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.554414Z digest=sha256:0592a03f7dae75b7909ec1eea7a5cc00cb1a07d7a4503af2d988da5e781ba20c

Observation 465a581a-dcfb-4462-9f82-f3ccd21ef63e · outbound

This paper cites Model evaluation for extreme risks.

LLM Cyber Evaluations Don't Capture Real-World Risk Model evaluation for extreme risks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.559543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.559543Z digest=sha256:9eb0eeb77221870316b6ab0b1c2f5858496e682317713cb86ffd2aa46de7e637

Observation 5329ec5e-6b9a-45b2-8b17-6366992c2269 · outbound

This paper cites Hacking CTFs with Plain Agents.

LLM Cyber Evaluations Don't Capture Real-World Risk Hacking CTFs with Plain Agents

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.564625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.564625Z digest=sha256:9d6e588b0183c7504fbe62bb05722f35db5aeab35a2935848f7e7e275e57c4b8

Observation a33ab2e4-6974-4332-9a83-d6bd7a00f22f · outbound

This paper cites Advanced AI evaluations may update, 2024.

LLM Cyber Evaluations Don't Capture Real-World Risk Advanced AI evaluations may update, 2024

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.286541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.570205Z digest=sha256:b041422c9e69ca3b333ef3f6454d7f14b7267a709fe6fb06e5baf18635ead2a0

Observation 7fe49fbc-4484-401c-a2a3-7794bab2e757 · outbound

This paper cites The near-term impact of AI on the cyber threat.

LLM Cyber Evaluations Don't Capture Real-World Risk The near-term impact of AI on the cyber threat

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.267306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.575782Z digest=sha256:d063a489c57fbd622e9c02ec9d5797a464db219d18f69f7ef262c724a938bae8

Observation 5ef8df1f-46c0-45de-876c-9a1ce7f36373 · outbound

This paper cites CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models.

LLM Cyber Evaluations Don't Capture Real-World Risk CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.581400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.581400Z digest=sha256:a6da27684e618b92741091ff16c153b1cc8b1b8468707f236be31f325332a3f6

Observation 405ac15e-b490-4ee3-996c-fbf0126befe2 · outbound

This paper cites Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models.

LLM Cyber Evaluations Don't Capture Real-World Risk Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.587373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.587373Z digest=sha256:67111551ff0a4ad83c893f77e0b15b43590dc5d8179ac79a1311b84cc5edf78a

Observation a7b32975-a92c-48b2-a216-75f3a2c4a383 · outbound

This paper cites write newline.

LLM Cyber Evaluations Don't Capture Real-World Risk write newline

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.593990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.593990Z digest=sha256:6f9577f54c836e07c4eb130a7a2eaeece39b8bc827ebbc06fdbad70ee7eecec3

Pith citing papers

Observation cb5c6c48-7cbe-4570-96c3-612dbd8bfe28 · inbound

Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks cites this paper.

Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks LLM Cyber Evaluations Don't Capture Real-World Risk

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T23:10:11.025662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T23:10:11.025662Z digest=sha256:f0c69d639fe7be1afc8df0c1b1abe45f12b7022ed6700409ea610c95f349c536

Observation 2064aa91-2129-48b3-b60f-39c01393918b · inbound

On the Surprising Efficacy of LLMs for Penetration-Testing cites this paper.

On the Surprising Efficacy of LLMs for Penetration-Testing LLM Cyber Evaluations Don't Capture Real-World Risk

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T21:10:06.577088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:10:06.577088Z digest=sha256:6f479b22b0c7de56243fa015d6289cbd2a758e29d096601465150d95c13c8f88

Observation 6374f621-c707-4497-b01a-f9b1c90a2c30 · inbound

The coordination gap in frontier AI safety policies cites this paper.

The coordination gap in frontier AI safety policies LLM Cyber Evaluations Don't Capture Real-World Risk

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-21T12:24:10.760233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T12:21:06.900491Z digest=sha256:68e9353b6b505476c5cf546e0954be106754ce4621e7753dc13aa327d373b88e

Observation dd764713-2be5-4852-b2a9-32a041fabb52 · inbound

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting cites this paper.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting LLM Cyber Evaluations Don't Capture Real-World Risk

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-07-03T02:07:33.283723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:f6bf4c7c32e7b83cd849e4625ecdbc599a2723f21647ff2c583c40c92653e19f