Pith. sign in

Paper Citation Record · LEDGER

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark

As of 17 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2608.11469.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11469 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:19:09.555981Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact6
  • verified fuzzy6
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b1eb8ad7-90b9-4218-b1a5-c45a73908c59 · outbound

This paper cites Vulnerability Detection with Code Language Models: How Far Are We?.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark Vulnerability Detection with Code Language Models: How Far Are We?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T14:19:09.463443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:19:09.463443Z digest=sha256:6dc3f8771716d25b93e47339285c0e9705cf2fec94830918c2a5563430e05140

Observation b91c5c13-c24f-4c0f-b0f9-44e36b867e0b · outbound

This paper cites Livecodebench: Holistic and contamination free evalua- tion of large language models for code.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark Livecodebench: Holistic and contamination free evalua- tion of large language models for code

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:19:10.117159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:19:09.482962Z digest=sha256:a0760884d8842301ac529efb9bac5fcfe219d2f9a42472212e2efae05e5e3871

Observation add1ed68-6e4b-4d14-a6cd-5e7bcb1d6cca · outbound

This paper cites REStack: A Large-Scale Dataset of Reverse Engineering Discussions from Stack Exchange.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark REStack: A Large-Scale Dataset of Reverse Engineering Discussions from Stack Exchange

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-15T14:19:09.995528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:19:09.488937Z digest=sha256:89d0e6b8604dcb0dc940f595ec07ced93c4626bc7789349f002b0c1d915069ec

Observation 89e9e2b8-9263-4ed2-ba8a-f2fbeccc797a · outbound

This paper cites REFORGE: A Method for Benchmarking LLMs' Reverse Engineering Capabilities in Decompiled Binary Function Naming.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark REFORGE: A Method for Benchmarking LLMs' Reverse Engineering Capabilities in Decompiled Binary Function Naming

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-15T14:19:09.966622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:19:09.494438Z digest=sha256:40ccf142e35c35e480c3fe7044696c3da7f842c3ddc77b658572ce5a8dd889e9

Observation fd51eebc-b6e7-4fdf-b5c8-a32a1b2c1bfc · outbound

This paper cites SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks?.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks?

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-15T14:19:09.932406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:19:09.500035Z digest=sha256:03c6745954c8ed1b9405a58e51c1db9cc6488976c7c32930a8ebdfd7c257e402

Observation c552c635-48c1-4632-b541-da7b525265cf · outbound

This paper cites ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T14:19:09.506467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:19:09.506467Z digest=sha256:dfec848f7ff4fa1335cde10064407bdcde229df88f7ff86789a74a06d7236a4a

Observation 3dafe5fe-d4a7-48a8-a715-060c674c64e5 · outbound

This paper cites VulDetectBench: Evaluating the Deep Capability of Vulnerability Detection with Large Language Models.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark VulDetectBench: Evaluating the Deep Capability of Vulnerability Detection with Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T14:19:09.511903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:19:09.511903Z digest=sha256:d00508cf769b2f58864b35ad2138ea5a96fd65d21c3ec011a9b18bc9f14d3de0

Observation 2385ced4-2f48-4612-9641-74581e10eab3 · outbound

This paper cites Patch-to-poc: A systematic study of agentic llm systems for linux kernel n-day reproduction.arXiv preprint arXiv:2602.07287,.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark Patch-to-poc: A systematic study of agentic llm systems for linux kernel n-day reproduction.arXiv preprint arXiv:2602.07287,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T14:19:09.517108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:19:09.517108Z digest=sha256:359b3631bb2ed2278540a2a77b86f9bbd16efa62765346097238358c20f4d62e

Observation 33f44ef8-e248-460e-82d0-050a37b5e4c2 · outbound

This paper cites ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T14:19:09.533769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:19:09.533769Z digest=sha256:76bfdb8e8cf7bd0e62a89207a296c4026b08b81b523c22452e1077dd5f34e6a0

Observation 59ea09b2-ab52-4a31-920b-e7bd3b2cc4d7 · outbound

This paper cites REBENCH: A Procedural, Fair-by-Construction Benchmark for LLMs on Stripped-Binary Types and Names (Extended Version).

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark REBENCH: A Procedural, Fair-by-Construction Benchmark for LLMs on Stripped-Binary Types and Names (Extended Version)

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-15T14:19:09.718401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:19:09.538836Z digest=sha256:a975f9e7bb3ebe0a77ac706d3488c7dd3aec10f1e3aaef085a8d0c0e5c58c1ed

Observation fbfa9057-47ef-4d9a-9aec-a9d5f361fe27 · outbound

This paper cites Cybench: A framework for evaluating cyber- security capabilities and risks of language models.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark Cybench: A framework for evaluating cyber- security capabilities and risks of language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:19:10.084545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:19:09.544423Z digest=sha256:d22136f4488a6d175a38833b10e612282df05453b8c0cd42c8e4c75d805876bb

Observation 868f2da6-28b3-4df1-bd39-6bb51e0ea382 · outbound

This paper cites CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T14:19:09.549877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:19:09.549877Z digest=sha256:503ead2fe9cabcb198a71c5418eb7525220f1ab63612799fd43cd03cd95e0da6

Observation 735dac0f-0924-4810-95f8-6c070051ae83 · outbound

This paper cites Training language model agents to find vulnerabilities with ctf-dojo.arXiv preprint arXiv:2508.18370,.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark Training language model agents to find vulnerabilities with ctf-dojo.arXiv preprint arXiv:2508.18370,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T14:19:09.555981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:19:09.555981Z digest=sha256:b9f99cb41b273fbda5e0d57c5e0a5c4cb6026ef324fa06ebf8dbd76ee7bf2f87

Observation 1aa92109-06f1-4f88-88dd-7ba8b870c5d6 · outbound

This paper cites CREBench: Evaluating Large Language Models in Cryptographic Binary Reverse Engineering.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark CREBench: Evaluating Large Language Models in Cryptographic Binary Reverse Engineering

Reference 1994

Resolution
verified exact
local_arxiv, observed 2026-08-15T14:19:10.063822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:19:09.451651Z digest=sha256:65a162d974c6c2cc76291a3cb43b92ba646e5aafac10824c296250fb457dcd5e

Observation 1efe0ece-f99c-458b-bf16-0dbf85826279 · outbound

This paper cites The concept assignment problem in program understanding.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark The concept assignment problem in program understanding

Reference 2005

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:19:10.175322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:19:09.441883Z digest=sha256:11347e8bdc2aeb92ee89b653240bdd3a7dbf316f41c38e416b2dfdeca9fa3609

Observation c75fbb3c-ea5c-49f6-8c25-ae639e94bf46 · outbound

This paper cites Benchmarking binary type inference techniques in decompilers.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark Benchmarking binary type inference techniques in decompilers

Reference 2008

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:19:10.101141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:19:09.521963Z digest=sha256:37bca616b85f70359e7a2f702288792b711d335c76c1c9a6d5f5b4ff4b9a1d75

Observation 2292012d-bf99-4722-addb-1cc959017dc0 · outbound

This paper cites De- compilebench: A comprehensive benchmark for evaluating decompilers in real-world scenarios.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark De- compilebench: A comprehensive benchmark for evaluating decompilers in real-world scenarios

Reference 2011

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:19:10.154010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:19:09.470783Z digest=sha256:d27825ce502f1b17618543ea286f9ae0023b1f16b182d944583a0f53361a611a

Observation 34d9a0fe-8c3f-4e3e-9226-5a02dc55cb1d · outbound

This paper cites Cy- bergym: Evaluating ai agents’ real-world cybersecurity capabilities at scale.arXiv preprint arXiv:2506.02548,.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark Cy- bergym: Evaluating ai agents’ real-world cybersecurity capabilities at scale.arXiv preprint arXiv:2506.02548,

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-15T14:19:09.528096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:19:09.528096Z digest=sha256:0ac29921db77eea43a9fa480ea2afff9f6f7015f4ce95f5df70ba235a19f3742

Observation 8574ff04-1778-4338-95ee-1823f150f486 · outbound

This paper cites Look what you made us patch: 2025 zero-days in review.https://cloud.google.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark Look what you made us patch: 2025 zero-days in review.https://cloud.google

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:19:10.134531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:19:09.476129Z digest=sha256:11dd93f76dfb161781b72c5df582d4428328b5be16a66c2f4a682659c3bc7784

Observation 919065b6-03ce-4b70-968e-669227f06a61 · outbound

This paper cites CrackMeBench: Binary Reverse Engineering for Agents.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark CrackMeBench: Binary Reverse Engineering for Agents

Reference 2026

Resolution
verified exact
local_arxiv, observed 2026-08-15T14:19:10.035977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:19:09.458239Z digest=sha256:b5acca3dd0ce50669bce9f404841caff6c7eb3a1223fa4356f62b865c562a957

Pith citing papers

No inbound Pith citation observations are available.