Pith. sign in

Paper Citation Record · LEDGER

Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

As of 3 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 39 inbound Pith citation observations for arXiv:2408.08926.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.08926 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 39 of 39 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T00:49:47.662275Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T18:37:31.168670Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0feb9130-5989-413f-818a-86b3d31b83ba · inbound

AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents cites this paper.

AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T01:35:51.135418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-14T01:35:50.992477Z digest=sha256:396be1791974c128a42d3b80848ab190426fc0c2486463192553081514b0d83d

Observation 936f24ff-bbee-4b68-afa7-b98308ff62ee · inbound

Frontier Models are Capable of In-context Scheming cites this paper.

Frontier Models are Capable of In-context Scheming Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T14:22:01.635191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-16T14:22:01.616448Z digest=sha256:fa89ad7dd48affdcb0b5752fda1246988b9e2bd3b61d98b4c477c919df597336

Observation be75ed09-53c6-4815-a5d7-86f23bc03ac8 · inbound

Humanity's Last Exam cites this paper.

Humanity's Last Exam Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.350304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:9dfad7581af4fd05bc2da43f02f5aa6a358551851d9558025960b131029744ca

Observation 00418eaf-ba6c-4c13-b005-b4182b51cdd1 · inbound

ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation cites this paper.

ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T04:42:04.771929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T04:37:33.942379Z digest=sha256:d13886a6176187fcbcbd6a7306f422d00a69b718b1fdade369040d9428b29dc9

Observation d5932e34-fc2d-4149-bfda-417ff6e36dd9 · inbound

Quantifying Frontier LLM Capabilities for Container Sandbox Escape cites this paper.

Quantifying Frontier LLM Capabilities for Container Sandbox Escape Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-02T19:43:40.853518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:43:40.853518Z digest=sha256:669ff1cc4c22aff26004e5dcae22d71e4502c3830c468ac2fc3f3eceabe966d7

Observation d0189277-b3d4-40eb-b3dc-9def2a74622a · inbound

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing cites this paper.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 132

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:45:52.939979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:ff938bdae6a1e6ff8eb552bdccc4049c45144fdde1c59e9fd750f3ac7bda2498

Observation 3c7fb686-9676-4ddb-b92a-76b5e3a3191a · inbound

SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills cites this paper.

SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T16:40:51.515451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:40:51.515451Z digest=sha256:2c86e08317b80d64619960d86cd51e8e46ab3a7492ad366c067a53458703a009

Observation 58f31584-7ea9-41de-ab10-7890120d3c43 · inbound

AlphaEval: Evaluating Agents in Production cites this paper.

AlphaEval: Evaluating Agents in Production Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:45:58.817569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:30:51.886471Z digest=sha256:44dd6f42a9474f2a53c46f45148983babd15ede34a6ca9219b700dafa0ae0b41

Observation 670e3a77-f49b-4e6a-b904-6acda6aaf242 · inbound

Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks cites this paper.

Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:06:19.087529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T06:02:32.399075Z digest=sha256:534f5079ac818834cce22156c556f735e9051c6b0954bf3e58699cbeb9025844

Observation 98a4d09b-b3c1-483c-8fd2-b68f5d949d6f · inbound

Can LLMs be Effective Code Contributors? A Study on Open-source Projects cites this paper.

Can LLMs be Effective Code Contributors? A Study on Open-source Projects Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:46:10.852344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T08:09:33.211692Z digest=sha256:1ecd7d5204f41169360a900a307b15d9daf971acd1bb6a60058b0403339a4561

Observation af10d1ef-b6f1-4213-8cdb-52e0b21494d8 · inbound

Dynamic Cyber Ranges cites this paper.

Dynamic Cyber Ranges Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T22:16:53.348658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T03:04:03.611481Z digest=sha256:584c648537ab652780898c74d59650b821e084f7dda9ca3eb744c323f6411429

Observation 901fe0bd-f957-4028-ad60-c19de5e945c4 · inbound

Risk Reporting for Developers' Internal AI Model Use cites this paper.

Risk Reporting for Developers' Internal AI Model Use Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:16:16.399595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T17:47:21.321820Z digest=sha256:47a3f13175f0d31e6b9728f42796a8a3a68ca7b44f95e5368720916254ca11ea

Observation f3479441-714c-411a-b0ed-5fb2f8e5e57e · inbound

XekRung Technical Report cites this paper.

XekRung Technical Report Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:51:14.588143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-09T20:55:10.400291Z digest=sha256:e0c786fc48d368fa32df63cd459f16f508ea57d5e200c9f9a279eadeeb6ec722

Observation 6f2e4b70-ed75-4997-8ad4-4ded7d60ee4d · inbound

Trace: Unmasking AI Attack Agents Through Terminal Behavior Fingerprinting cites this paper.

Trace: Unmasking AI Attack Agents Through Terminal Behavior Fingerprinting Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:41:18.994757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-09T15:18:43.426681Z digest=sha256:8fb5b8918b2ed0303c502bb8696b38b1598e82db854ce2427f00af3fd5453d87

Observation 8c487224-e811-4da9-89bd-d92904e4ba7f · inbound

Autonomous Adversary: Red-Teaming in the age of LLM cites this paper.

Autonomous Adversary: Red-Teaming in the age of LLM Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:26:10.508553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T09:09:58.566171Z digest=sha256:4e28490f82fb3d6b595126c1d2b1d9eafedaae2b8880ea9795f1ed1ab5a4acef

Observation fb629ce5-673b-4c01-a4d6-b413708b2c71 · inbound

Patch2Vuln: Agentic Reconstruction of Vulnerabilities from Linux Distribution Binary Patches cites this paper.

Patch2Vuln: Agentic Reconstruction of Vulnerabilities from Linux Distribution Binary Patches Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:31:11.685007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T08:51:33.112454Z digest=sha256:22236c7c2d8473998aaef70a5320811c7e135998456d8b8e480ab25fb0b2c734

Observation 056324e2-fb99-4104-b553-afcf4b86cac3 · inbound

CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios cites this paper.

CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:10:57.348437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T01:54:46.391345Z digest=sha256:fbe4d0b7e0cbd6536bf720b45ac25254bb8b37598008fb015844df3ad8fbb1a1

Observation f204818c-98b3-493f-ab84-07171811d024 · inbound

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World cites this paper.

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:16:27.972213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:34:55.538935Z digest=sha256:5793cb1bf7d2f7d708fe1aa664365937c6c934a8ad6610e71361e42210d6c9d7

Observation 24ac8b9d-0e14-4b24-9694-01757c5bc2be · inbound

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World cites this paper.

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T14:22:27.655588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:22:27.655588Z digest=sha256:8319a5126e361f1434b23c1a0ec306da4826d8d21a7f7bca5b4b7528d6c76378

Observation 83f64bae-f5e2-4773-88c2-c67e3df1b3d2 · inbound

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications cites this paper.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-19T23:32:52.563721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:a22afe3d817dc9ac1c28bb9e3179d14fa101d4dffcbb01520228b3b7dfbf6941

Observation 33d9d0cc-241a-4c13-8df6-342a25d52c98 · inbound

Benchmarking Mythos-Linked Bug Rediscovery cites this paper.

Benchmarking Mythos-Linked Bug Rediscovery Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T23:17:57.552731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T23:14:34.164854Z digest=sha256:ba281c3210c1dc5faa5b509ab02c141a631303df4e44ab91735abec53158939f

Observation 8584ade9-95b6-446c-bf9e-18902bde487f · inbound

DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows cites this paper.

DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T10:18:11.749362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T10:16:38.920528Z digest=sha256:41b259c738c6d51257a7780bf500c56300d9120007ae112fe90a46b7b2099c20

Observation 305c9254-29ad-43f3-8569-8b25cfa8cddf · inbound

HIDBench: Benchmarking Large Language Models for Host-Based Intrusion Detection cites this paper.

HIDBench: Benchmarking Large Language Models for Host-Based Intrusion Detection Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T08:54:45.823514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T08:52:34.079804Z digest=sha256:de9c43f407daa796143ee15b498776b840174a8db5058e54a69aa2002892b342

Observation 7e042dcc-a9f5-4faf-8701-a82bb982bb3f · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:07.929199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T05:50:28.114140Z digest=sha256:bb2ce98ecae110bfef0835a36ca1b2dd2e9f7abe6675e934d62d9e5a1e5e5aa5

Observation 001041d1-9921-45e9-9866-f3db5b16d6ed · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:06:43.185520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-25T06:05:27.736494Z digest=sha256:23170cb80a7f97d9153ccf7f5873a5ae5620e3e47b898344bd84a43e501a1acc

Observation 0e1e3c4c-a3a3-4fc0-85fc-a7b7e84f2e31 · inbound

Cybersecurity AI (CAI) Dataset cites this paper.

Cybersecurity AI (CAI) Dataset Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:13:26.856759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:07:49.656453Z digest=sha256:0d88e18f4fdfd194662e8de1989f506282996f06987ad82d40ea3308b2e90608

Observation ab106e19-42cb-47ed-b501-4f3b97b30ebf · inbound

Stateful Online Monitoring Catches Distributed Agent Attacks cites this paper.

Stateful Online Monitoring Catches Distributed Agent Attacks Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:56:11.064697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-28T21:54:44.072929Z digest=sha256:2b7a0663984f11f8266c82f29a245781077bee1dcbd14c30fec11d83d110e494

Observation d56a6cf9-90bf-410a-b13a-6f3bd59e972b · inbound

An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios cites this paper.

An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:48:46.344299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-27T03:39:36.657903Z digest=sha256:7ba100abe8b243f62f79321c9aca9cb25b81193ff61288bcf3502e78563029e4

Observation 99c16f9a-1b02-49bb-b4e3-14ce6bf99e4a · inbound

Poisoned Playbooks: Demystifying Knowledge Poisoning Effects on AI Security Agents cites this paper.

Poisoned Playbooks: Demystifying Knowledge Poisoning Effects on AI Security Agents Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T17:40:01.102907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-25T23:27:14.610097Z digest=sha256:ee00188d830e084f172c385614663757deaf2653a518853c0c69a2d56c1efce5

Observation f1e2d4fa-4aef-4e74-9e82-9d19555f84d9 · inbound

Direct Causation in International Humanitarian Law and the Challenge of AI-Mediated Civilian Cyber Operations cites this paper.

Direct Causation in International Humanitarian Law and the Challenge of AI-Mediated Civilian Cyber Operations Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:54:22.125038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T07:49:28.819402Z digest=sha256:01e14a1c2c5f9e7fc832331a823bc9b4a52e5812f7c3aea76c5b9dc697b86591

Observation b35eadfd-1caa-44b8-a2e7-2eb5d55a9303 · inbound

Mastermind: Strategy-grounded Learning for Repository-Scale Vulnerability Reproduction cites this paper.

Mastermind: Strategy-grounded Learning for Repository-Scale Vulnerability Reproduction Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:08:21.502047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T14:02:06.173081Z digest=sha256:79292473fef0036b3f9be744564eaf8efcf1cba11751a3f75697d75a64e2398f

Observation ef5f186d-bd5e-4b54-b7d1-5f3115183312 · inbound

ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents cites this paper.

ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-10T18:37:31.169937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-10T18:29:50.731038Z digest=sha256:7010f611ca022431694278a92105386ae65cff7b83d97fc4f713aa7f46a2f20e

Observation 65c92e11-d7bc-443a-8ac8-4bfc8b6376c6 · inbound

ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents cites this paper.

ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T00:49:47.662275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:49:47.662275Z digest=sha256:9f4eb9bf2f67ddad899b795cdf31c2094b050ea6a0dfa859abd00038ff2e9fb3

Observation e5ec7e9b-7c8e-48ae-9d35-a983f0fcfb48 · inbound

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents cites this paper.

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T23:44:37.323064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:44:37.323064Z digest=sha256:6ac72fd5fbf8320572588d8c9faed8afc1ce10c0dfaa15bfbef1d57e84cace6e

Observation ceff4a75-aac0-40df-b438-27b6d540f7d0 · inbound

Harmonizing AI Safety Thresholds cites this paper.

Harmonizing AI Safety Thresholds Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T21:24:15.880141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:24:15.880141Z digest=sha256:4bf358950a28e713ffed8d284aa355a2686c12e2dba11b8c9a1825434d8bf148

Observation 23646153-646d-483d-acfc-45130e22460a · inbound

RECEIPT: Deterministic, Reward-Hacking-Resistant Verification for White-Box Agentic XSS Discovery cites this paper.

RECEIPT: Deterministic, Reward-Hacking-Resistant Verification for White-Box Agentic XSS Discovery Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T15:04:09.259606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:04:09.259606Z digest=sha256:b0e7e4a1024260629c22deb622124af90b8b3a0a9074fb51ff18e2179a5ede24

Observation 911f64eb-a401-452b-86ca-d0fc8217fff4 · inbound

Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks cites this paper.

Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T06:49:26.953693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:49:26.953693Z digest=sha256:41440ae734b21c760d3a944c88af37cf665daa3f51516767be0062b173a592b0

Observation 579da816-ab2d-4ffb-98f7-0939b6df73a4 · inbound

The Disruptive Impact of Large Language Models on Capture the Flag Competitions and the Path Toward Fair Play cites this paper.

The Disruptive Impact of Large Language Models on Capture the Flag Competitions and the Path Toward Fair Play Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T02:31:26.073853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:31:26.073853Z digest=sha256:5fa5eef3b748c17b759222ae9ec4f7da38bd1b1f4c08a9aa1e77acf5263ce18f

Observation 8c75942d-277b-4329-a925-2a8d7343f2de · inbound

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents cites this paper.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:37.291705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:37.291705Z digest=sha256:deea31812072e58ce6bb48163b4cd55ddedde558d289a393e49d3732f61af0d3