Pith. sign in

Paper Citation Record · LEDGER

LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2312.12575.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.12575 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:01:16.105652Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T10:45:43.061376Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 15390e56-20ed-4974-80f4-af460d0c14df · inbound

XOXO: Stealthy Cross-Origin Context Poisoning Attacks against AI Coding Assistants cites this paper.

XOXO: Stealthy Cross-Origin Context Poisoning Attacks against AI Coding Assistants LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:57:17.086140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T23:55:14.887237Z digest=sha256:f26e2c92aa4d6520608f4ada327d1087aa315a25a3aaa968823e5a26fe169cc3

Observation a4f1c9de-b850-44b7-9fb2-c4ceaaf81f90 · inbound

Mono: Is Your "Clean" Vulnerability Dataset Really Solvable? Exposing and Trapping Undecidable Patches and Beyond cites this paper.

Mono: Is Your "Clean" Vulnerability Dataset Really Solvable? Exposing and Trapping Undecidable Patches and Beyond LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:01:16.105652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:01:16.105652Z digest=sha256:6feeadcba9cd28497b06aad472965af2341794a40791a0c608e2e3940e8f155e

Observation f450ee49-39b8-4ee6-82a8-6824a1947f21 · inbound

Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges cites this paper.

Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:05.983497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:35:05.983497Z digest=sha256:0983eff8177a42945af5389114bcf0e773ea9242b450bfe86771395757aa31a8

Observation 3ba98a77-e4f3-4c0b-8aeb-54fdb87972f6 · inbound

Large Language Models in Cybersecurity: Applications, Vulnerabilities, and Defense Techniques cites this paper.

Large Language Models in Cybersecurity: Applications, Vulnerabilities, and Defense Techniques LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:28.904690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:28.904690Z digest=sha256:f159fae188e5c04c8fec1a157c664bd421f2095e2d3eff9367a29c9b76843f3a

Observation 466ba651-e18b-424e-bade-bc8ad19b0068 · inbound

Rethinking LLM-Based RTL Code Optimization Via Timing Logic Metamorphosis cites this paper.

Rethinking LLM-Based RTL Code Optimization Via Timing Logic Metamorphosis LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:07:20.536909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:07:20.536909Z digest=sha256:0b4ab69f7d7c172a6ff12e46044a7bb22945d8814524ea24c81ab6ad4b6ce653

Observation 58e5ca18-62e4-43de-ba8b-c3b41903be0e · inbound

LogicEval: A Systematic Framework for Evaluating Automated Repair Techniques for Logical Vulnerabilities in Real-World Software cites this paper.

LogicEval: A Systematic Framework for Evaluating Automated Repair Techniques for Logical Vulnerabilities in Real-World Software LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:11:03.926724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:07:44.158768Z digest=sha256:28d0ed9c7c5e81aa296a03c2ef795616c99f80224739b666acff5d4f4b160717

Observation 99891289-6814-44ea-83bb-a4ada86260e3 · inbound

ASSEMBLAGE-DEEPHISTORY: A Cross-Build Binary Dataset with Temporal Coverage cites this paper.

ASSEMBLAGE-DEEPHISTORY: A Cross-Build Binary Dataset with Temporal Coverage LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 78

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T09:41:21.750390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T09:40:21.746161Z digest=sha256:58b49259ce4aa040bd71e09eb439f6d1b800635c14d123a414d6d75b559cf467

Observation e280e68f-1216-4b7d-9d5b-f34c9e5e8d36 · inbound

An Empirical Study of Security Calibration in Large Language Models for Code cites this paper.

An Empirical Study of Security Calibration in Large Language Models for Code LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:45:43.063064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:04:41.223743Z digest=sha256:a08f54cb2778453be4485d827cd14ce9e1d2b2454a391b638ff776c98e317b1d