{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:LLPZ5DSAWR5YZT5NCZYVN5IMBV","short_pith_number":"pith:LLPZ5DSA","schema_version":"1.0","canonical_sha256":"5adf9e8e40b47b8ccfad167156f50c0d41b658de8bcd4ffc04c5f05102cd8a57","source":{"kind":"arxiv","id":"2408.08926","version":4},"attestation_state":"computed","paper":{"title":"Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI","cs.CL","cs.CY","cs.LG"],"primary_cat":"cs.CR","authors_text":"Andy K. Zhang, Ari Glenn, Celeste Menders, Dan Boneh, Daniel E. Ho, Daniel Zamoshchin, Derek Askaryar, Donovan Jasper, Eliot Jones, Gashon Hussein, Gautham Raghupathi, Joey Ji, Justin W. Lin, Kenny Osele, Leo Glikbarg, Mike Yang, Nathan Tran, Neil Perry, Percy Liang, Polycarpos Yiorkadjis, Pura Peetathawatchai, Rinnara Sangpisit, Rishi Alluri, Riya Dulepet, Samantha Liu, Teddy Zhang, Vikram Sivashankar","submitted_at":"2024-08-15T17:23:10Z","abstract_excerpt":"Language Model (LM) agents for cybersecurity that are capable of autonomously identifying vulnerabilities and executing exploits have potential to cause real-world impact. Policymakers, model providers, and researchers in the AI and cybersecurity communities are interested in quantifying the capabilities of such agents to help mitigate cyberrisk and investigate opportunities for penetration testing. Toward that end, we introduce Cybench, a framework for specifying cybersecurity tasks and evaluating agents on those tasks. We include 40 professional-level Capture the Flag (CTF) tasks from 4 dist"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2408.08926","kind":"arxiv","version":4},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.CR","submitted_at":"2024-08-15T17:23:10Z","cross_cats_sorted":["cs.AI","cs.CL","cs.CY","cs.LG"],"title_canon_sha256":"9d606e70cdbe558916d80b5c5b48be2c362c9b90458bc0f487d5242334657ae4","abstract_canon_sha256":"61cff41d13a23379bab61317532c9e9b970345c4ba5ac04fe2f3242726f1da04"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T10:48:18.488909Z","signature_b64":"cah+xz+FBUq7cPjtH+0/MOy4+JaqPJcLpAtY/qQRLjqoDxYTt0CLXOK2gxdKT7klzH9DGMfU5ZMKVb6NJmABCQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"5adf9e8e40b47b8ccfad167156f50c0d41b658de8bcd4ffc04c5f05102cd8a57","last_reissued_at":"2026-07-05T10:48:18.488389Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T10:48:18.488389Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI","cs.CL","cs.CY","cs.LG"],"primary_cat":"cs.CR","authors_text":"Andy K. Zhang, Ari Glenn, Celeste Menders, Dan Boneh, Daniel E. Ho, Daniel Zamoshchin, Derek Askaryar, Donovan Jasper, Eliot Jones, Gashon Hussein, Gautham Raghupathi, Joey Ji, Justin W. Lin, Kenny Osele, Leo Glikbarg, Mike Yang, Nathan Tran, Neil Perry, Percy Liang, Polycarpos Yiorkadjis, Pura Peetathawatchai, Rinnara Sangpisit, Rishi Alluri, Riya Dulepet, Samantha Liu, Teddy Zhang, Vikram Sivashankar","submitted_at":"2024-08-15T17:23:10Z","abstract_excerpt":"Language Model (LM) agents for cybersecurity that are capable of autonomously identifying vulnerabilities and executing exploits have potential to cause real-world impact. Policymakers, model providers, and researchers in the AI and cybersecurity communities are interested in quantifying the capabilities of such agents to help mitigate cyberrisk and investigate opportunities for penetration testing. Toward that end, we introduce Cybench, a framework for specifying cybersecurity tasks and evaluating agents on those tasks. We include 40 professional-level Capture the Flag (CTF) tasks from 4 dist"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2408.08926","kind":"arxiv","version":4},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2408.08926/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2408.08926","created_at":"2026-07-05T10:48:18.488454+00:00"},{"alias_kind":"arxiv_version","alias_value":"2408.08926v4","created_at":"2026-07-05T10:48:18.488454+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2408.08926","created_at":"2026-07-05T10:48:18.488454+00:00"},{"alias_kind":"pith_short_12","alias_value":"LLPZ5DSAWR5Y","created_at":"2026-07-05T10:48:18.488454+00:00"},{"alias_kind":"pith_short_16","alias_value":"LLPZ5DSAWR5YZT5N","created_at":"2026-07-05T10:48:18.488454+00:00"},{"alias_kind":"pith_short_8","alias_value":"LLPZ5DSA","created_at":"2026-07-05T10:48:18.488454+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":29,"internal_anchor_count":1,"sample":[{"citing_arxiv_id":"2607.07774","citing_title":"ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents","ref_index":11,"is_internal_anchor":true},{"citing_arxiv_id":"2606.24402","citing_title":"Poisoned Playbooks: Demystifying Knowledge Poisoning Effects on AI Security Agents","ref_index":52,"is_internal_anchor":false},{"citing_arxiv_id":"2606.17114","citing_title":"An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios","ref_index":12,"is_internal_anchor":false},{"citing_arxiv_id":"2607.01764","citing_title":"Mastermind: Strategy-grounded Learning for Repository-Scale Vulnerability Reproduction","ref_index":31,"is_internal_anchor":false},{"citing_arxiv_id":"2605.31593","citing_title":"Stateful Online Monitoring Catches Distributed Agent Attacks","ref_index":32,"is_internal_anchor":false},{"citing_arxiv_id":"2606.29175","citing_title":"Direct Causation in International Humanitarian Law and the Challenge of AI-Mediated Civilian Cyber Operations","ref_index":12,"is_internal_anchor":false},{"citing_arxiv_id":"2605.28146","citing_title":"Cybersecurity AI (CAI) Dataset","ref_index":22,"is_internal_anchor":false},{"citing_arxiv_id":"2605.22643","citing_title":"Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety","ref_index":94,"is_internal_anchor":false},{"citing_arxiv_id":"2605.21773","citing_title":"HIDBench: Benchmarking Large Language Models for Host-Based Intrusion Detection","ref_index":5,"is_internal_anchor":false},{"citing_arxiv_id":"2605.22643","citing_title":"Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety","ref_index":94,"is_internal_anchor":false},{"citing_arxiv_id":"2605.19099","citing_title":"DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows","ref_index":55,"is_internal_anchor":false},{"citing_arxiv_id":"2605.17413","citing_title":"Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications","ref_index":41,"is_internal_anchor":false},{"citing_arxiv_id":"2605.17416","citing_title":"Benchmarking Mythos-Linked Bug Rediscovery","ref_index":42,"is_internal_anchor":false},{"citing_arxiv_id":"2507.14201","citing_title":"ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation","ref_index":60,"is_internal_anchor":false},{"citing_arxiv_id":"2412.04984","citing_title":"Frontier Models are Capable of In-context Scheming","ref_index":37,"is_internal_anchor":false},{"citing_arxiv_id":"2410.09024","citing_title":"AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents","ref_index":34,"is_internal_anchor":false},{"citing_arxiv_id":"2605.10834","citing_title":"From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World","ref_index":50,"is_internal_anchor":false},{"citing_arxiv_id":"2604.24966","citing_title":"Risk Reporting for Developers' Internal AI Model Use","ref_index":49,"is_internal_anchor":false},{"citing_arxiv_id":"2604.24184","citing_title":"Dynamic Cyber Ranges","ref_index":11,"is_internal_anchor":false},{"citing_arxiv_id":"2604.23340","citing_title":"Can LLMs be Effective Code Contributors? A Study on Open-source Projects","ref_index":29,"is_internal_anchor":false},{"citing_arxiv_id":"2605.06601","citing_title":"Patch2Vuln: Agentic Reconstruction of Vulnerabilities from Linux Distribution Binary Patches","ref_index":38,"is_internal_anchor":false},{"citing_arxiv_id":"2605.06486","citing_title":"Autonomous Adversary: Red-Teaming in the age of LLM","ref_index":1,"is_internal_anchor":false},{"citing_arxiv_id":"2605.01186","citing_title":"Trace: Unmasking AI Attack Agents Through Terminal Behavior Fingerprinting","ref_index":30,"is_internal_anchor":false},{"citing_arxiv_id":"2605.00072","citing_title":"XekRung Technical Report","ref_index":58,"is_internal_anchor":false},{"citing_arxiv_id":"2604.12162","citing_title":"AlphaEval: Evaluating Agents in Production","ref_index":13,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/LLPZ5DSAWR5YZT5NCZYVN5IMBV","json":"https://pith.science/pith/LLPZ5DSAWR5YZT5NCZYVN5IMBV.json","graph_json":"https://pith.science/api/pith-number/LLPZ5DSAWR5YZT5NCZYVN5IMBV/graph.json","events_json":"https://pith.science/api/pith-number/LLPZ5DSAWR5YZT5NCZYVN5IMBV/events.json","paper":"https://pith.science/paper/LLPZ5DSA"},"agent_actions":{"view_html":"https://pith.science/pith/LLPZ5DSAWR5YZT5NCZYVN5IMBV","download_json":"https://pith.science/pith/LLPZ5DSAWR5YZT5NCZYVN5IMBV.json","view_paper":"https://pith.science/paper/LLPZ5DSA","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2408.08926&json=true","fetch_graph":"https://pith.science/api/pith-number/LLPZ5DSAWR5YZT5NCZYVN5IMBV/graph.json","fetch_events":"https://pith.science/api/pith-number/LLPZ5DSAWR5YZT5NCZYVN5IMBV/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/LLPZ5DSAWR5YZT5NCZYVN5IMBV/action/timestamp_anchor","attest_storage":"https://pith.science/pith/LLPZ5DSAWR5YZT5NCZYVN5IMBV/action/storage_attestation","attest_author":"https://pith.science/pith/LLPZ5DSAWR5YZT5NCZYVN5IMBV/action/author_attestation","sign_citation":"https://pith.science/pith/LLPZ5DSAWR5YZT5NCZYVN5IMBV/action/citation_signature","submit_replication":"https://pith.science/pith/LLPZ5DSAWR5YZT5NCZYVN5IMBV/action/replication_record"}},"created_at":"2026-07-05T10:48:18.488454+00:00","updated_at":"2026-07-05T10:48:18.488454+00:00"}