Pith. sign in

Paper Citation Record · LEDGER

Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 63 inbound Pith citation observations for arXiv:2408.08926.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.08926 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 63 of 63 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:58:10.496145Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T18:37:31.168670Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0feb9130-5989-413f-818a-86b3d31b83ba · inbound

AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents cites this paper.

AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T01:35:51.135418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-14T01:35:50.992477Z digest=sha256:cdda1572876faf23b671b35a1676250e082bc4913a219e365e17c5cb2ff8467c

Observation 0bcc9c81-ae4d-4420-877f-98c379cde089 · inbound

Safety case template for frontier AI: A cyber inability argument cites this paper.

Safety case template for frontier AI: A cyber inability argument Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T22:01:11.592200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T22:01:11.592200Z digest=sha256:b54736779f94ede693234519b8e14fee8b808db868fe3c1dd4018b9b13f7474b

Observation 5ba35a28-52d5-4eba-a12e-10d9ef4b35e0 · inbound

Evaluating Large Language Models' Capability to Launch Fully Automated Spear Phishing Campaigns: Validated on Human Subjects cites this paper.

Evaluating Large Language Models' Capability to Launch Fully Automated Spear Phishing Campaigns: Validated on Human Subjects Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:17.065290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:20:17.065290Z digest=sha256:59de7ec010b7ddb92532654f61e36456cce381919e0c31d14c730eb3460e3423

Observation e9e62439-b645-43eb-bb71-dc29c3d57d92 · inbound

HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing cites this paper.

HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T01:00:00.190563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T01:00:00.190563Z digest=sha256:c0a4aa1d57cea3d33a5b6fde46e698bead40b7da683b7d44d643f63126dcbd2f

Observation 936f24ff-bbee-4b68-afa7-b98308ff62ee · inbound

Frontier Models are Capable of In-context Scheming cites this paper.

Frontier Models are Capable of In-context Scheming Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T14:22:01.635191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-16T14:22:01.616448Z digest=sha256:49c7c4760e5b3e4617785ad92330efbd6f179faa8ed71f807f2b1541bf6982ed

Observation be75ed09-53c6-4815-a5d7-86f23bc03ac8 · inbound

Humanity's Last Exam cites this paper.

Humanity's Last Exam Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.350304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:1b8d42482286017cedd181900336820a14fb21167eca96b2106b0993b78f72fd

Observation 405ac15e-b490-4ee3-996c-fbf0126befe2 · inbound

LLM Cyber Evaluations Don't Capture Real-World Risk cites this paper.

LLM Cyber Evaluations Don't Capture Real-World Risk Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.587373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.587373Z digest=sha256:11810d0ca2861bdfded52c1a31bc8db639cd22f093675df44047c0524e5c8cd0

Observation fcd94e89-ae74-4576-a430-c68e8cb1942c · inbound

Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks cites this paper.

Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T23:10:11.149366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T23:10:11.149366Z digest=sha256:bcc412d47a7cd3c228de12e3c0b078942966848216fd8005fab64cffe6aa79d2

Observation 4dd76008-2c7c-4b8e-a7d3-32227712fbce · inbound

A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management cites this paper.

A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:44.828234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:48:44.828234Z digest=sha256:8ed19a9b7b1d9d79bb04eed4cf36cb5db7f8044e24232d67872a6c88ba0e1b4d

Observation 572ba4b6-6ed9-42b1-973a-eed5a1ec2da0 · inbound

LLMs unlock new paths to monetizing exploits cites this paper.

LLMs unlock new paths to monetizing exploits Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T20:58:10.496145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:58:10.496145Z digest=sha256:54ec518aa591f23343fdfd7b624b985861a2828002e03df759b1ad5c3791adea

Observation 9c21991c-eadb-487d-9a76-7927a1c1cf7b · inbound

CRAKEN: Cybersecurity LLM Agent with Knowledge-Based Execution cites this paper.

CRAKEN: Cybersecurity LLM Agent with Knowledge-Based Execution Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:52.914666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:23:52.914666Z digest=sha256:15c2b972f99c723f663e873719f5278ef5c4eefe5d6ba588e3480c1551cac78d

Observation bc7718ec-e1a9-48e5-92f1-5b9dca55f357 · inbound

Improving LLM Agents with Reinforcement Learning on Cryptographic CTF Challenges cites this paper.

Improving LLM Agents with Reinforcement Learning on Cryptographic CTF Challenges Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:35.406709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:35.406709Z digest=sha256:9dd41c43ebb6c14930faef9873e6b0eb47dc04c8afb9e94f94012e6ee719db52

Observation 7794c2eb-86ab-446d-9979-d9af2c171628 · inbound

Recognition Without Mitigation: Ethical Frameworks in Autonomous Offensive-LLM Agent Research cites this paper.

Recognition Without Mitigation: Ethical Frameworks in Autonomous Offensive-LLM Agent Research Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:10:20.166338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:10:20.166338Z digest=sha256:c79c4c9cda189dc2b6d176b441e6f1fc68ea0c4570b702b1731d448b50bfb8cf

Observation 71a0b358-a58c-4eae-a135-ff47b069821d · inbound

From Promise to Peril: Rethinking Cybersecurity Red and Blue Teaming in the Age of LLMs cites this paper.

From Promise to Peril: Rethinking Cybersecurity Red and Blue Teaming in the Age of LLMs Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T00:35:03.624759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:35:03.624759Z digest=sha256:7bdb75c74da52493d4ad62aa4fb9f118070b589ae8d8a95eab062031edf1368a

Observation cb576ed1-53fe-4da0-ab9c-9c13e0aa1f7e · inbound

Establishing Best Practices for Building Rigorous Agentic Benchmarks cites this paper.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:28.287235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:28.287235Z digest=sha256:5aa8b5518a39645bb991ea2f0b096ad2b5309a4e207fa9280463e1b1bb6936f6

Observation 290af98a-b2cf-432e-bfdb-cf3e875068bd · inbound

Evaluating the Critical Risks of Amazon's Nova Premier under the Frontier Model Safety Framework cites this paper.

Evaluating the Critical Risks of Amazon's Nova Premier under the Frontier Model Safety Framework Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:39:59.356503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:39:59.356503Z digest=sha256:50198ba02e2f4e25cd42f140fc680f794dd5b729792753c237525c4ec240a0bc

Observation fc9251f0-e653-4768-ba17-b7bf3c19b967 · inbound

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights cites this paper.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:34.325643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:55:34.325643Z digest=sha256:673ead8e23fcf3b134f3e84de4b7546af3b9caf7e28d3ef132bf35323d922e69

Observation 00418eaf-ba6c-4c13-b005-b4182b51cdd1 · inbound

ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation cites this paper.

ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T04:42:04.771929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-19T04:37:33.942379Z digest=sha256:0a221b8de8e60ae42c47d0fc5052db1f5eef9ec17280876cbd516bc26197ae89

Observation ca2ca46d-83e1-4bf8-b634-cbf1ce3a42f3 · inbound

Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report cites this paper.

Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T15:14:23.006196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:14:23.006196Z digest=sha256:58cc1361a0694e863594423f808124b401a4713897095babfd3edf8e6a33cc5d

Observation a2e33d96-9c84-4793-a4d4-08cedb088a61 · inbound

Agent Identity Evals: Measuring Agentic Identity cites this paper.

Agent Identity Evals: Measuring Agentic Identity Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.491612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.491612Z digest=sha256:a9b24415e3b372387513b8af9e38f47e1f11fd6e762ca7bdfee57f11a2919df5

Observation 5207f360-b370-4439-a14b-e9273761fc7f · inbound

Evaluation and Benchmarking of LLM Agents: A Survey cites this paper.

Evaluation and Benchmarking of LLM Agents: A Survey Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 114

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.858357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.858357Z digest=sha256:2faa3a42361f2568e194ef04e4a953a4300fa831a937c03195b31abd632dba78

Observation 063639a0-1e72-4497-a8a7-3174e901d29a · inbound

Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report cites this paper.

Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T05:57:29.522003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:57:29.522003Z digest=sha256:b69d19ff13184a812addc627b1178b3de64ff64e277f244efcf2edbd49985102

Observation 99d6da7f-8a57-4eda-a3e2-58af7100db2c · inbound

PoCo: Agentic Proof-of-Concept Exploit Generation for Smart Contracts cites this paper.

PoCo: Agentic Proof-of-Concept Exploit Generation for Smart Contracts Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T00:09:36.085164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:09:36.085164Z digest=sha256:e23781cc2ce5b2cfe32a321a6a225cd4ea5e648f631f8c0bced664cad5ae762d

Observation d5932e34-fc2d-4149-bfda-417ff6e36dd9 · inbound

Quantifying Frontier LLM Capabilities for Container Sandbox Escape cites this paper.

Quantifying Frontier LLM Capabilities for Container Sandbox Escape Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-02T19:43:40.853518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:43:40.853518Z digest=sha256:a637e727480b89dcf6e2969efdef3dd62f25dd4c36c9afa4f598d39fffda028c

Observation 8ffd15fe-f1cb-47df-870a-e1f7d5365307 · inbound

Quantifying Frontier LLM Capabilities for Container Sandbox Escape cites this paper.

Quantifying Frontier LLM Capabilities for Container Sandbox Escape Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T06:01:01.313606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:01:01.313606Z digest=sha256:6298839e89243aef9b81fc3fa90c85dbb00a8c194852bdfc5110c83d5e8870d8

Observation d0189277-b3d4-40eb-b3dc-9def2a74622a · inbound

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing cites this paper.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 132

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:45:52.939979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:f75bc1d0a5f04d5d93bab60146211a535b95880ccfb14be12e8cb81c45dbbacb

Observation 3c7fb686-9676-4ddb-b92a-76b5e3a3191a · inbound

SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills cites this paper.

SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T16:40:51.515451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:40:51.515451Z digest=sha256:402ce991cfbde0eff5eada8a6f7ea203195ed86477448a34d9429c8ff4698041

Observation 58f31584-7ea9-41de-ab10-7890120d3c43 · inbound

AlphaEval: Evaluating Agents in Production cites this paper.

AlphaEval: Evaluating Agents in Production Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:45:58.817569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T16:30:51.886471Z digest=sha256:dc7f013b92d0ca43d5be203e8a1b7de4d7f66867bcbf52586b319db795dacd44

Observation 670e3a77-f49b-4e6a-b904-6acda6aaf242 · inbound

Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks cites this paper.

Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:06:19.087529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T06:02:32.399075Z digest=sha256:0cfbe9ebf57dfc0f8268ba33855eae051cac12bf141aa76a2b85c207ff009519

Observation 98a4d09b-b3c1-483c-8fd2-b68f5d949d6f · inbound

Can LLMs be Effective Code Contributors? A Study on Open-source Projects cites this paper.

Can LLMs be Effective Code Contributors? A Study on Open-source Projects Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:46:10.852344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T08:09:33.211692Z digest=sha256:68b293bd77d644c4f065c9b9d0ac092002e6340fd30be953674db4ed3f6b5e0a

Observation af10d1ef-b6f1-4213-8cdb-52e0b21494d8 · inbound

Dynamic Cyber Ranges cites this paper.

Dynamic Cyber Ranges Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T22:16:53.348658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T03:04:03.611481Z digest=sha256:bdc60832712f50cc79279aa8ecdcdf13bfd49e599eac6deca2e737ce1873c7ad

Observation 901fe0bd-f957-4028-ad60-c19de5e945c4 · inbound

Risk Reporting for Developers' Internal AI Model Use cites this paper.

Risk Reporting for Developers' Internal AI Model Use Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:16:16.399595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-07T17:47:21.321820Z digest=sha256:cc187e44be9ce148fdb3864f95cf7491987cf4e8ef2ef9e3d8a12ef2abd40339

Observation f3479441-714c-411a-b0ed-5fb2f8e5e57e · inbound

XekRung Technical Report cites this paper.

XekRung Technical Report Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:51:14.588143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-09T20:55:10.400291Z digest=sha256:6b378dc1f24ba83a4656f60a4d9b35768743e9448aaae3e5761a6fb3035aa52c

Observation 6f2e4b70-ed75-4997-8ad4-4ded7d60ee4d · inbound

Trace: Unmasking AI Attack Agents Through Terminal Behavior Fingerprinting cites this paper.

Trace: Unmasking AI Attack Agents Through Terminal Behavior Fingerprinting Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:41:18.994757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-09T15:18:43.426681Z digest=sha256:14b452e0a14a0def44d396f70477d01a2a5f19a6e36ee688957b2a1d0e708bcf

Observation 8c487224-e811-4da9-89bd-d92904e4ba7f · inbound

Autonomous Adversary: Red-Teaming in the age of LLM cites this paper.

Autonomous Adversary: Red-Teaming in the age of LLM Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:26:10.508553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T09:09:58.566171Z digest=sha256:bb2e9165699c0c094112149a8ddbe16fcda3ae6daed66f3ab4171059d8e12f3f

Observation fb629ce5-673b-4c01-a4d6-b413708b2c71 · inbound

Patch2Vuln: Agentic Reconstruction of Vulnerabilities from Linux Distribution Binary Patches cites this paper.

Patch2Vuln: Agentic Reconstruction of Vulnerabilities from Linux Distribution Binary Patches Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:31:11.685007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T08:51:33.112454Z digest=sha256:76db08ec9f94b017b2876ced0150cd3568f605f02404a5fbf2e42f21111245e6

Observation 056324e2-fb99-4104-b553-afcf4b86cac3 · inbound

CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios cites this paper.

CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:10:57.348437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T01:54:46.391345Z digest=sha256:e98eea52bdf1418c434de68b2c6ded15e954a24c94f5385497dc7303bcdfe214

Observation f204818c-98b3-493f-ab84-07171811d024 · inbound

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World cites this paper.

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:16:27.972213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T03:34:55.538935Z digest=sha256:bf750091b4d4a4f1f74fd1b48b4269a5287d8b32b3411e2976986bc8f160db10

Observation 24ac8b9d-0e14-4b24-9694-01757c5bc2be · inbound

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World cites this paper.

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T14:22:27.655588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:22:27.655588Z digest=sha256:1c79aad458fcb7eb482451f624c5ab977708b53c68ad9a4ea8400a631513179f

Observation 83f64bae-f5e2-4773-88c2-c67e3df1b3d2 · inbound

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications cites this paper.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-19T23:32:52.563721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:933bffa69281c6930a028161c3e36eb8377983ccf92b5f8bf706377bbc272f35

Observation 33d9d0cc-241a-4c13-8df6-342a25d52c98 · inbound

Benchmarking Mythos-Linked Bug Rediscovery cites this paper.

Benchmarking Mythos-Linked Bug Rediscovery Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T23:17:57.552731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-19T23:14:34.164854Z digest=sha256:1181bd5c68edd03fbeb56f3675fbef35fb63fd14fa7a8e78a88582bf9d6af72d

Observation 8584ade9-95b6-446c-bf9e-18902bde487f · inbound

DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows cites this paper.

DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T10:18:11.749362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:16:38.920528Z digest=sha256:e550623f8caeb05ae68269af594351482d7965dec07c0d1d7fa46990ab9db946

Observation 305c9254-29ad-43f3-8569-8b25cfa8cddf · inbound

HIDBench: Benchmarking Large Language Models for Host-Based Intrusion Detection cites this paper.

HIDBench: Benchmarking Large Language Models for Host-Based Intrusion Detection Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T08:54:45.823514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T08:52:34.079804Z digest=sha256:e02202a27f4fc7eee0ebaa115914bf3d27c3dc0a9effe599b703f154663409c6

Observation 7e042dcc-a9f5-4faf-8701-a82bb982bb3f · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:07.929199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T05:50:28.114140Z digest=sha256:c47cdbe522da82d4c3b8996be0a8cfd4f6d8746cdb3d21e638bdcb8700d4f6ed

Observation 001041d1-9921-45e9-9866-f3db5b16d6ed · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:06:43.185520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-25T06:05:27.736494Z digest=sha256:5ab88a40a86e2e1b0b92c0a3a840cc2d6767387e5aa2ea0f6ef6df98f6afb520

Observation 0e1e3c4c-a3a3-4fc0-85fc-a7b7e84f2e31 · inbound

Cybersecurity AI (CAI) Dataset cites this paper.

Cybersecurity AI (CAI) Dataset Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:13:26.856759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T12:07:49.656453Z digest=sha256:d4a9ab75d81327d928e00b6ec8d1e883fe085d029b4ea243ef8fe019eb17fa13

Observation ab106e19-42cb-47ed-b501-4f3b97b30ebf · inbound

Stateful Online Monitoring Catches Distributed Agent Attacks cites this paper.

Stateful Online Monitoring Catches Distributed Agent Attacks Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:56:11.064697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T21:54:44.072929Z digest=sha256:d9e3a31cb0ab607aa59fad43a487094581a35b23803d4fe9349dce3d9f9a8330

Observation d56a6cf9-90bf-410a-b13a-6f3bd59e972b · inbound

An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios cites this paper.

An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:48:46.344299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T03:39:36.657903Z digest=sha256:43ca4d82966729613e69be7b94b9702f88090833bed97ea31c93fc6403b10ce6

Observation 99c16f9a-1b02-49bb-b4e3-14ce6bf99e4a · inbound

Poisoned Playbooks: Demystifying Knowledge Poisoning Effects on AI Security Agents cites this paper.

Poisoned Playbooks: Demystifying Knowledge Poisoning Effects on AI Security Agents Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T17:40:01.102907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-25T23:27:14.610097Z digest=sha256:a0d2ef0f1445e64fed82122d7e4da007dd493426fea42dc9639bc29d494ee60d

Observation f1e2d4fa-4aef-4e74-9e82-9d19555f84d9 · inbound

Direct Causation in International Humanitarian Law and the Challenge of AI-Mediated Civilian Cyber Operations cites this paper.

Direct Causation in International Humanitarian Law and the Challenge of AI-Mediated Civilian Cyber Operations Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:54:22.125038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T07:49:28.819402Z digest=sha256:03c4b7891524e7fd50d615ace6ae6f556962cc741a22fba1c7d2a0ba95a594d8

Observation b35eadfd-1caa-44b8-a2e7-2eb5d55a9303 · inbound

Mastermind: Strategy-grounded Learning for Repository-Scale Vulnerability Reproduction cites this paper.

Mastermind: Strategy-grounded Learning for Repository-Scale Vulnerability Reproduction Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:08:21.502047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-03T14:02:06.173081Z digest=sha256:7bd0067853517dfa5881f972f0b41d8c6f6ea31b36a11fa7be1e5993cc2b2ef3

Observation ef5f186d-bd5e-4b54-b7d1-5f3115183312 · inbound

ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents cites this paper.

ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-10T18:37:31.169937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T18:29:50.731038Z digest=sha256:59110c79990e3a0ed9279ca33c912944539c6549a8e5685b9002c2c703a08218

Observation 65c92e11-d7bc-443a-8ac8-4bfc8b6376c6 · inbound

ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents cites this paper.

ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T00:49:47.662275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:49:47.662275Z digest=sha256:5e61c6c14d3e84789c625cf8109cb43cec66286d329befe58f75961b785e5ff2

Observation e5ec7e9b-7c8e-48ae-9d35-a983f0fcfb48 · inbound

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents cites this paper.

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T23:44:37.323064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:44:37.323064Z digest=sha256:5bfb7d5f4f1efb686e398d195589bc25b8b77d7109736fa31c923a0bc54bf333

Observation ceff4a75-aac0-40df-b438-27b6d540f7d0 · inbound

Harmonizing AI Safety Thresholds cites this paper.

Harmonizing AI Safety Thresholds Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T21:24:15.880141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:24:15.880141Z digest=sha256:a07abf74768f356a08f9a462ed81cd06e066097476f6c5799ec225bf613265f1

Observation 23646153-646d-483d-acfc-45130e22460a · inbound

RECEIPT: Deterministic, Reward-Hacking-Resistant Verification for White-Box Agentic XSS Discovery cites this paper.

RECEIPT: Deterministic, Reward-Hacking-Resistant Verification for White-Box Agentic XSS Discovery Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T15:04:09.259606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:04:09.259606Z digest=sha256:b3bf64123d8f97069d62cd33492126e0e19ff48a472d86172d349e5acfbb0247

Observation 365145fe-aa15-4a18-8967-cceae3af76e2 · inbound

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction cites this paper.

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T15:36:02.669356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:36:02.669356Z digest=sha256:fb2cc536a8bf61e9bacbdcc34618effa7df388b8824653e9c6b8d44dc85d203e

Observation 911f64eb-a401-452b-86ca-d0fc8217fff4 · inbound

Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks cites this paper.

Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T06:49:26.953693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:49:26.953693Z digest=sha256:ee057fd469316c7b504d4e6c39f845423efd99a9b660a142539ac8da8aa53e30

Observation 579da816-ab2d-4ffb-98f7-0939b6df73a4 · inbound

The Disruptive Impact of Large Language Models on Capture the Flag Competitions and the Path Toward Fair Play cites this paper.

The Disruptive Impact of Large Language Models on Capture the Flag Competitions and the Path Toward Fair Play Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T02:31:26.073853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:31:26.073853Z digest=sha256:4aa2c16cd6d520ab65b0838ba851de01169780b68b49e56ceb8a5c1c1eef0161

Observation 8c75942d-277b-4329-a925-2a8d7343f2de · inbound

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents cites this paper.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:37.291705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:37.291705Z digest=sha256:8d1d7b4c453b2b2d04961ffa059472940d527db7cc30a8bbc0bdb165a3838c73

Observation 4a545ca8-c695-4c69-9201-fab9a77b2fae · inbound

ThreatForest: Multi-Agent Attack Tree Generation with Pluggable TTP Framework Mapping cites this paper.

ThreatForest: Multi-Agent Attack Tree Generation with Pluggable TTP Framework Mapping Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T15:27:18.744781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:27:18.744781Z digest=sha256:ef84ff65961dd3fe8d0294a3f686b9a6a01f8f139bf4b27d010aa34df255c3cf

Observation ccc8c6be-b185-4ddf-95ee-b317d5299044 · inbound

Tiny Enough to Break In: Agentic Remote Access Trojans Powered by Small Language Models cites this paper.

Tiny Enough to Break In: Agentic Remote Access Trojans Powered by Small Language Models Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T04:19:18.902915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:19:18.902915Z digest=sha256:91401449890323c2db7c57e826236f163f5368ea5da40d224abba346ff08ed10

Observation 221f4235-aa5c-479f-a954-1eb0a5e64326 · inbound

CyberForge: Verified Vulnerability Injection at Repository Level for Cybersecurity Agent Training cites this paper.

CyberForge: Verified Vulnerability Injection at Repository Level for Cybersecurity Agent Training Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:02.923242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:37:02.923242Z digest=sha256:d34b7cf211807b8b30575c166bcbde8521825fe87e8d9215c42a9d94613dcd9f