Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 3 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 39 inbound Pith citation observations for arXiv:2408.08926.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T00:49:47.662275Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T18:37:31.168670Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 0feb9130-5989-413f-818a-86b3d31b83ba · inbound
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 936f24ff-bbee-4b68-afa7-b98308ff62ee · inbound
Frontier Models are Capable of In-context Scheming Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation be75ed09-53c6-4815-a5d7-86f23bc03ac8 · inbound
Humanity's Last Exam Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 00418eaf-ba6c-4c13-b005-b4182b51cdd1 · inbound
ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d5932e34-fc2d-4149-bfda-417ff6e36dd9 · inbound
Quantifying Frontier LLM Capabilities for Container Sandbox Escape Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0189277-b3d4-40eb-b3dc-9def2a74622a · inbound
Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 132
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 3c7fb686-9676-4ddb-b92a-76b5e3a3191a · inbound
SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58f31584-7ea9-41de-ab10-7890120d3c43 · inbound
AlphaEval: Evaluating Agents in Production Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 670e3a77-f49b-4e6a-b904-6acda6aaf242 · inbound
Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 98a4d09b-b3c1-483c-8fd2-b68f5d949d6f · inbound
Can LLMs be Effective Code Contributors? A Study on Open-source Projects Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation af10d1ef-b6f1-4213-8cdb-52e0b21494d8 · inbound
Dynamic Cyber Ranges Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 901fe0bd-f957-4028-ad60-c19de5e945c4 · inbound
Risk Reporting for Developers' Internal AI Model Use Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f3479441-714c-411a-b0ed-5fb2f8e5e57e · inbound
XekRung Technical Report Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 6f2e4b70-ed75-4997-8ad4-4ded7d60ee4d · inbound
Trace: Unmasking AI Attack Agents Through Terminal Behavior Fingerprinting Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8c487224-e811-4da9-89bd-d92904e4ba7f · inbound
Autonomous Adversary: Red-Teaming in the age of LLM Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation fb629ce5-673b-4c01-a4d6-b413708b2c71 · inbound
Patch2Vuln: Agentic Reconstruction of Vulnerabilities from Linux Distribution Binary Patches Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 056324e2-fb99-4104-b553-afcf4b86cac3 · inbound
CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f204818c-98b3-493f-ab84-07171811d024 · inbound
From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 24ac8b9d-0e14-4b24-9694-01757c5bc2be · inbound
From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83f64bae-f5e2-4773-88c2-c67e3df1b3d2 · inbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 33d9d0cc-241a-4c13-8df6-342a25d52c98 · inbound
Benchmarking Mythos-Linked Bug Rediscovery Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8584ade9-95b6-446c-bf9e-18902bde487f · inbound
DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 305c9254-29ad-43f3-8569-8b25cfa8cddf · inbound
HIDBench: Benchmarking Large Language Models for Host-Based Intrusion Detection Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7e042dcc-a9f5-4faf-8701-a82bb982bb3f · inbound
Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 001041d1-9921-45e9-9866-f3db5b16d6ed · inbound
Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 0e1e3c4c-a3a3-4fc0-85fc-a7b7e84f2e31 · inbound
Cybersecurity AI (CAI) Dataset Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ab106e19-42cb-47ed-b501-4f3b97b30ebf · inbound
Stateful Online Monitoring Catches Distributed Agent Attacks Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d56a6cf9-90bf-410a-b13a-6f3bd59e972b · inbound
An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 99c16f9a-1b02-49bb-b4e3-14ce6bf99e4a · inbound
Poisoned Playbooks: Demystifying Knowledge Poisoning Effects on AI Security Agents Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f1e2d4fa-4aef-4e74-9e82-9d19555f84d9 · inbound
Direct Causation in International Humanitarian Law and the Challenge of AI-Mediated Civilian Cyber Operations Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b35eadfd-1caa-44b8-a2e7-2eb5d55a9303 · inbound
Mastermind: Strategy-grounded Learning for Repository-Scale Vulnerability Reproduction Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ef5f186d-bd5e-4b54-b7d1-5f3115183312 · inbound
ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 65c92e11-d7bc-443a-8ac8-4bfc8b6376c6 · inbound
ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5ec7e9b-7c8e-48ae-9d35-a983f0fcfb48 · inbound
Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ceff4a75-aac0-40df-b438-27b6d540f7d0 · inbound
Harmonizing AI Safety Thresholds Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23646153-646d-483d-acfc-45130e22460a · inbound
RECEIPT: Deterministic, Reward-Hacking-Resistant Verification for White-Box Agentic XSS Discovery Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 911f64eb-a401-452b-86ca-d0fc8217fff4 · inbound
Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 579da816-ab2d-4ffb-98f7-0939b6df73a4 · inbound
The Disruptive Impact of Large Language Models on Capture the Flag Competitions and the Path Toward Fair Play Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c75942d-277b-4329-a925-2a8d7343f2de · inbound
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.