Pith. sign in

Paper Citation Record · LEDGER

SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks?

As of 13 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2605.26548.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.26548 v2

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T13:10:41.870343Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6cd2488a-f694-4ced-898c-f6f4d4c8cbf8 · outbound

This paper cites SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?.

SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks? SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T13:10:39.062654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:10:39.062654Z digest=sha256:ee6d2883d893b343fe80da3fda9216bf4c5acd37aeea2e6da3d5831a71ca1b25

Observation 57c6980f-9f27-46d7-b3e5-6fb1eb5cb49b · outbound

This paper cites not on target page.

SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks? not on target page

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T13:10:41.870343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:10:41.870343Z digest=sha256:6247527011b8202cf8553273c551a38b97f7487ee09caf08b7fed9b776b7adb0

Observation bedb53be-0590-4fac-ab8d-c6e1fd5599ad · outbound

This paper cites an unresolved cited work.

SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks? Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T13:10:39.332264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:10:39.332264Z digest=sha256:5218d8c0c078768bc5c4f65ec9c6e75380d091fe4ab9ab0731544ae3534e0660

Observation 7b050198-709c-47a4-aef4-f42089960311 · outbound

This paper cites Haoyu Li, Xijia Che, Yanhao Wang, Xiaojing Liao, and Luyi Xing.

SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks? Haoyu Li, Xijia Che, Yanhao Wang, Xiaojing Liao, and Luyi Xing

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T13:10:39.556798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:10:39.556798Z digest=sha256:b89de8c48b6421f7d8552262903645a2de47312bed1e6e35356244f1d7a993ad

Observation 6ecd2d0d-a224-486d-a18a-db12c78eb25f · outbound

This paper cites A Dual-Loop Agent Framework for Automated Vulnerability Reproduction.arXiv preprint arXiv:2602.05721,.

SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks? A Dual-Loop Agent Framework for Automated Vulnerability Reproduction.arXiv preprint arXiv:2602.05721,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T13:10:39.720845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:10:39.720845Z digest=sha256:c2ec261bc13c9bb1807eaad9c3f999051a612b2256272a4b481e88e571fcedcb

Observation e5277e74-807e-4536-a468-587eb7e192b3 · outbound

This paper cites Automated Vulnerability Validation and Verification: A Large Language Model Approach.arXiv preprint arXiv:2509.24037,.

SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks? Automated Vulnerability Validation and Verification: A Large Language Model Approach.arXiv preprint arXiv:2509.24037,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T13:10:39.806055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:10:39.806055Z digest=sha256:acd25751ca7a8c7e704eefc17bd924d416a9f8f7499b49245e4af77662ece588

Observation 0a148fe4-ec80-4b6a-9cba-682de91936f7 · outbound

This paper cites ARVO: Atlas of Reproducible Vulnerabilities for Open-Source Software.

SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks? ARVO: Atlas of Reproducible Vulnerabilities for Open-Source Software

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T13:10:39.908889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:10:39.908889Z digest=sha256:314963a79432b78383590c3fdea5e5834c913e6d52efcb24fc97d9542c336a7d

Observation b97b4a4f-fc81-439c-b441-46569dc84bd0 · outbound

This paper cites Codex Security.https://help.openai.com/articles/20001107, 2026a.

SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks? Codex Security.https://help.openai.com/articles/20001107, 2026a

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T13:10:40.144901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:10:40.144901Z digest=sha256:62097dd6e249ed411ed6c695fa54b9768071787bd89865ffc9b9d7ecff978bb6

Observation 8314cb41-31ea-4907-946f-eb1ec8b816b6 · outbound

This paper cites Patch-to-PoC: A Systematic Study of Agentic LLM Systems for Linux Kernel N-Day Reproduction.arXiv preprint arXiv:2602.07287,.

SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks? Patch-to-PoC: A Systematic Study of Agentic LLM Systems for Linux Kernel N-Day Reproduction.arXiv preprint arXiv:2602.07287,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T13:10:40.242592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:10:40.242592Z digest=sha256:5454604ddf1bc3659c2ce204d68d3610e6ae338c08a7609def153ba1cb78f517

Observation ca79777c-087b-4274-9059-94f8acdc096b · outbound

This paper cites To Err is Machine: Vulnerability Detection Challenges LLM Reasoning.

SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks? To Err is Machine: Vulnerability Detection Challenges LLM Reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T13:10:40.467725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:10:40.467725Z digest=sha256:066368ea3ac44a787d5e0ae6c984c7e19f16b0ca4c60088ff56bfbc3f537a5e4

Observation 4dbca145-14d2-462f-976e-139eadd99315 · outbound

This paper cites From Naptime to Big Sleep: Using Large Language Models To Catch Vulnerabilities In Real-World Code.https://googleprojectzero.blogspot.com/2024/ 10/from-naptime-to-big-sleep.html,.

SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks? From Naptime to Big Sleep: Using Large Language Models To Catch Vulnerabilities In Real-World Code.https://googleprojectzero.blogspot.com/2024/ 10/from-naptime-to-big-sleep.html,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T13:10:40.602706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:10:40.602706Z digest=sha256:d7106d679f62fdc33c607d71bab0860bd73f668aff09071a65831a2091f2d4a9

Observation d49d6713-073a-4bee-8f1a-bf369c216c04 · outbound

This paper cites Coskun, and Gianluca Stringhini.

SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks? Coskun, and Gianluca Stringhini

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T13:10:40.822958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:10:40.822958Z digest=sha256:60399e644f68a71825bff9155f29bb2469c16b1534e113653859eda30127b2df

Observation 11ee100e-dcd4-4f1c-8d52-2f76fa6c5f7f · outbound

This paper cites From CVE Entries to Verifiable Exploits: An Automated Multi-Agent Framework for Reproducing CVEs.arXiv preprint arXiv:2509.01835,.

SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks? From CVE Entries to Verifiable Exploits: An Automated Multi-Agent Framework for Reproducing CVEs.arXiv preprint arXiv:2509.01835,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T13:10:40.977062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:10:40.977062Z digest=sha256:ef20cd5415e7c4b37651050ce6765e940855c5b71af2056eb302e8e99b1a465a

Observation f6040990-466e-48c2-9882-f4293c255258 · outbound

This paper cites Zichao Wei, Jun Zeng, Ming Wen, Zeliang Yu, Kai Cheng, Yiding Zhu, Jingyi Guo, Shiqi Zhou, Le Yin, Xiaodong Su, and Zhechao Ma.

SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks? Zichao Wei, Jun Zeng, Ming Wen, Zeliang Yu, Kai Cheng, Yiding Zhu, Jingyi Guo, Shiqi Zhou, Le Yin, Xiaodong Su, and Zhechao Ma

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T13:10:41.130975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:10:41.130975Z digest=sha256:5ea77ebe1d67c77031258125c33914432bcddd153a4251f492cf2aa0220ae07c

Observation 4101701b-1883-46c1-ab19-33749dc1f3b6 · outbound

This paper cites Why Is CSP Failing? Trends and Challenges in CSP Adoption.

SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks? Why Is CSP Failing? Trends and Challenges in CSP Adoption

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T13:10:41.278384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:10:41.278384Z digest=sha256:1c3cb2be59640fb160cc856a1c08827b79e4b0abd220938bf6b9ddac3b100c32

Observation 8877253a-ba49-450b-be39-ee7688061346 · outbound

This paper cites ProgramBench: Can Language Models Rebuild Programs From Scratch?.

SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks? ProgramBench: Can Language Models Rebuild Programs From Scratch?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T13:10:41.424053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:10:41.424053Z digest=sha256:48ee57befd84e560396398a70f0e2e6f26609a60ea7293bf280cbcf4cedf1264

Observation 5df73039-f2a9-4449-9229-9451da11a600 · outbound

This paper cites Bhatia, Vikram Sivashankar, Yuxuan Bao, Dawn Song, Dan Boneh, Daniel Ho, and Percy Liang.

SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks? Bhatia, Vikram Sivashankar, Yuxuan Bao, Dawn Song, Dan Boneh, Daniel Ho, and Percy Liang

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T13:10:41.589836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:10:41.589836Z digest=sha256:6c324794623f872fdd1148074e1524d933e368e7064a344e9e9cb0181d440b54

Observation 11cfe1db-ee4d-4bb6-ac90-edee42727ef3 · outbound

This paper cites A Systematic Study on Generating Web Vulnerability Proof-of-Concepts Using Large Language Models.arXiv preprint arXiv:2510.10148,.

SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks? A Systematic Study on Generating Web Vulnerability Proof-of-Concepts Using Large Language Models.arXiv preprint arXiv:2510.10148,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T13:10:41.698626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:10:41.698626Z digest=sha256:2eaa1284ec37fbd473e5b44f4bbb150961749838f2398270e2a4315089d2fd53

Observation a058625f-3062-4926-af18-8be7bda83573 · outbound

This paper cites AnyPoC: Universal Proof-of-Concept Test Generation for Scalable LLM-Based Bug Detection.

SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks? AnyPoC: Universal Proof-of-Concept Test Generation for Scalable LLM-Based Bug Detection

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T13:10:41.788247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:10:41.788247Z digest=sha256:8196b67c86a085d97afa24580a1b6c3a6d837d747d55011474d254428ddbe6c7

Observation 5f986cd1-df99-4d0d-b9d3-14d0fac539bd · outbound

This paper cites FaultLine: Automated Proof-of-Vulnerability Generation Using LLM Agents.

SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks? FaultLine: Automated Proof-of-Vulnerability Generation Using LLM Agents

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-02T13:10:40.014171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:10:40.014171Z digest=sha256:e18e3b9375a0c4278576c365ad7be7a54ce09755683bf1cb1792d940a72a2a8b

Observation 94eda813-e819-4078-9c09-76ab404abcdb · outbound

This paper cites ZeroDayBench: Evalu- ating LLM Agents on Unseen Zero-Day Vulnerabilities for Cyberdefense.arXiv preprint arXiv:2603.02297,.

SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks? ZeroDayBench: Evalu- ating LLM Agents on Unseen Zero-Day Vulnerabilities for Cyberdefense.arXiv preprint arXiv:2603.02297,

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-02T13:10:39.440957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:10:39.440957Z digest=sha256:be61cff2e75a0140b78f867a6eea3ef33a38a8ff44d8159ec991027769eb3ce5

Observation 4dfba7bd-f1fb-4ba8-ba56-651694406fa0 · outbound

This paper cites PoCGen: Generating Proof-of-Concept Exploits for Vulnerabilities in Npm Packages.

SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks? PoCGen: Generating Proof-of-Concept Exploits for Vulnerabilities in Npm Packages

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T13:10:40.346309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:10:40.346309Z digest=sha256:971f5f445aeff6b54dc29029158c4c4d2b069b2bd7cfab614b160ac14a43231e

Observation 41e5fc11-617f-4f8b-b02b-316fb5fafc8a · outbound

This paper cites Context Length Alone Hurts LLM Performance Despite Perfect Retrieval.

SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks? Context Length Alone Hurts LLM Performance Despite Perfect Retrieval

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T13:10:39.179091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:10:39.179091Z digest=sha256:e8d6784c0d9b4a792c6444756a7c4a04356ae964457c1407547067f8723f3e54

Observation a7d3e999-f2db-44ae-b16f-49e938e61a34 · outbound

This paper cites Introducing Claude Opus 4.6.

SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks? Introducing Claude Opus 4.6

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T13:10:38.958167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:10:38.958167Z digest=sha256:bd952049c5f39ac9022ab3769e57dbdac533a19d6024d2e7879eaa244f38ce0e

Pith citing papers

No inbound Pith citation observations are available.