Pith. sign in

Paper Citation Record · LEDGER

LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2312.12575.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.12575 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:50:24.807578Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T10:45:43.061376Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7259b0ee-5486-44fc-91f9-2b4a95b60943 · inbound

Combining GPT and Code-Based Similarity Checking for Effective Smart Contract Vulnerability Detection cites this paper.

Combining GPT and Code-Based Similarity Checking for Effective Smart Contract Vulnerability Detection LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T04:59:42.493802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:59:42.493802Z digest=sha256:846744060aecad6cae6521efb5d86ed57fd6b3b1c730a138af3fff0f632ab2f6

Observation 8ad7b8cf-605d-492d-86b9-08d4f961b81a · inbound

Integrating Artificial Open Generative Artificial Intelligence into Software Supply Chain Security cites this paper.

Integrating Artificial Open Generative Artificial Intelligence into Software Supply Chain Security LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T01:02:43.645372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:02:43.645372Z digest=sha256:af905aa710aabe9ee7616c09f0b983d23b3434d1bfce06ad24f5bcd44146f02d

Observation 15390e56-20ed-4974-80f4-af460d0c14df · inbound

XOXO: Stealthy Cross-Origin Context Poisoning Attacks against AI Coding Assistants cites this paper.

XOXO: Stealthy Cross-Origin Context Poisoning Attacks against AI Coding Assistants LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:57:17.086140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T23:55:14.887237Z digest=sha256:e7c66c39de7e864c0c75475744ed390fe4882076ad871d64dd568adbcc336390

Observation 4887733e-19e5-480c-a0a7-0c63c7da7da9 · inbound

Automatically Generating Rules of Malicious Software Packages via Large Language Model cites this paper.

Automatically Generating Rules of Malicious Software Packages via Large Language Model LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:24.807578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:24.807578Z digest=sha256:545cd57b71a38004eff198ea4decdddd3510df7458347b50fb43697f2581de0a

Observation a4f1c9de-b850-44b7-9fb2-c4ceaaf81f90 · inbound

Mono: Is Your "Clean" Vulnerability Dataset Really Solvable? Exposing and Trapping Undecidable Patches and Beyond cites this paper.

Mono: Is Your "Clean" Vulnerability Dataset Really Solvable? Exposing and Trapping Undecidable Patches and Beyond LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:01:16.105652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:01:16.105652Z digest=sha256:6b77d4b92d38f9d3061c9648a4b6097ae962d781265b4cbf2e9bae37b2b06590

Observation f450ee49-39b8-4ee6-82a8-6824a1947f21 · inbound

Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges cites this paper.

Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:05.983497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:35:05.983497Z digest=sha256:73ca6e083d6e293b1eebff837819c69553aa3a6a2d1f725d28b2df8651d68449

Observation 3ba98a77-e4f3-4c0b-8aeb-54fdb87972f6 · inbound

Large Language Models in Cybersecurity: Applications, Vulnerabilities, and Defense Techniques cites this paper.

Large Language Models in Cybersecurity: Applications, Vulnerabilities, and Defense Techniques LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:28.904690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:28.904690Z digest=sha256:f79f8b8c90a2bfd0dd8f6e84ecbb9c025fe74a54ccde39c29ff386bd183e3f89

Observation 466ba651-e18b-424e-bade-bc8ad19b0068 · inbound

Rethinking LLM-Based RTL Code Optimization Via Timing Logic Metamorphosis cites this paper.

Rethinking LLM-Based RTL Code Optimization Via Timing Logic Metamorphosis LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:07:20.536909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:07:20.536909Z digest=sha256:1775c9676b9dde4f8a6515cd5d480461489955b704fe0f1fbf0188a34ad35cd3

Observation 58e5ca18-62e4-43de-ba8b-c3b41903be0e · inbound

LogicEval: A Systematic Framework for Evaluating Automated Repair Techniques for Logical Vulnerabilities in Real-World Software cites this paper.

LogicEval: A Systematic Framework for Evaluating Automated Repair Techniques for Logical Vulnerabilities in Real-World Software LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:11:03.926724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T15:07:44.158768Z digest=sha256:6e4186c71e4254579f60f94c146f567bb64092e9742445f267afd9f37d2ecc82

Observation 99891289-6814-44ea-83bb-a4ada86260e3 · inbound

ASSEMBLAGE-DEEPHISTORY: A Cross-Build Binary Dataset with Temporal Coverage cites this paper.

ASSEMBLAGE-DEEPHISTORY: A Cross-Build Binary Dataset with Temporal Coverage LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 78

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T09:41:21.750390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T09:40:21.746161Z digest=sha256:e9ded2a65c61f0ec3e012931391f21c236dca981e3bdb495fbeac49120507e8d

Observation e280e68f-1216-4b7d-9d5b-f34c9e5e8d36 · inbound

An Empirical Study of Security Calibration in Large Language Models for Code cites this paper.

An Empirical Study of Security Calibration in Large Language Models for Code LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:45:43.063064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-01T05:04:41.223743Z digest=sha256:1a6da309a7d4e2f8821a4512bfd5877f379fc50aec72f6d88fb8ffaec134de3c

Observation 9bb6f895-ca7d-4ae9-876e-c048d34df059 · inbound

Activation Probes Surface Code-Security Signals that the Model's Output Misses cites this paper.

Activation Probes Surface Code-Security Signals that the Model's Output Misses LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T13:35:03.919947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:35:03.919947Z digest=sha256:cd15cb2594de644c29ca57bc2c12f20e9fdb42779b7a33dd72eb4358ca956525