Pith. sign in

Paper Citation Record · LEDGER

Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 45 inbound Pith citation observations for arXiv:2312.04724.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.04724 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 45 of 45 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:33:32.868576Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

18
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d71ed716-aa49-42ae-a6c4-491f44c60ad0 · inbound

StarCoder 2 and The Stack v2: The Next Generation cites this paper.

StarCoder 2 and The Stack v2: The Next Generation Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 164

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T17:28:22.675474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T17:28:22.353355Z digest=sha256:f75557cdd519767d7a00bd63e0eb7674816a6bb6af0700b4f155de1050d50578

Observation 4299a779-96e4-4419-9105-def9b99382a0 · inbound

Precision or Peril: A PoC of Python Code Quality from Quantized Large Language Models cites this paper.

Precision or Peril: A PoC of Python Code Quality from Quantized Large Language Models Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:33:14.955539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T17:30:54.204300Z digest=sha256:65269f905b56ad8bd288b9d7d4068cddaa22160d4592729650d2064e75ccb95d

Observation 8ff9165e-9e70-4832-be1e-67814bb5b3d5 · inbound

CWEval: Outcome-driven Evaluation on Functionality and Security of LLM Code Generation cites this paper.

CWEval: Outcome-driven Evaluation on Functionality and Security of LLM Code Generation Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:32.868576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:32.868576Z digest=sha256:813bee36d6a105ad2ff6251b770ac9b8a58ced8ec3611651855390b562dad6cb

Observation f303afb9-da19-4460-9313-d870f3e455bd · inbound

Humanity's Last Exam cites this paper.

Humanity's Last Exam Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.368040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:0d5d9fa018cfc3b04d4403d1ff52e513afd00505b428e899e45f18b05f196785

Observation 301c7b7c-519f-472b-9145-8a07f717f1db · inbound

LLMSecConfig: An LLM-Based Approach for Fixing Software Container Misconfigurations cites this paper.

LLMSecConfig: An LLM-Based Approach for Fixing Software Container Misconfigurations Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T13:45:55.122362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:45:55.122362Z digest=sha256:4ca32c152bdbd916b93666851ffd488881cfab742bbe82412ef7a885bfd9c73d

Observation 3e89f275-13dc-49a2-8358-52d8b34ab118 · inbound

Benchmarking Prompt Engineering Techniques for Secure Code Generation with GPT Models cites this paper.

Benchmarking Prompt Engineering Techniques for Secure Code Generation with GPT Models Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T17:00:30.464049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:00:30.464049Z digest=sha256:269201e684a00c39284ee94593bb26dff80b3377dbb7fe095616c5a5b3a11541

Observation 259fba11-9ee5-4183-a195-facc5c07f03b · inbound

Training Language Models to Generate Quality Code with Program Analysis Feedback cites this paper.

Training Language Models to Generate Quality Code with Program Analysis Feedback Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:04:58.231629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:04:58.231629Z digest=sha256:b029aa08282a961f9a9e7a6c63e05d855ef5fef0b8372b8dc54fdad6503b2967

Observation 6f606523-bef9-46e5-b411-10bf0bb1bb23 · inbound

MCP Safety Training: Learning to Refuse Falsely Benign MCP Exploits using Improved Preference Alignment cites this paper.

MCP Safety Training: Learning to Refuse Falsely Benign MCP Exploits using Improved Preference Alignment Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:10.267103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:10.267103Z digest=sha256:abe6e3fe277be37ce9b8b40102ed2a6c2a717d29371c8cd5940efa54c7791e04

Observation 2a740e82-3223-4dfb-9470-ccff15569fbc · inbound

Developing a Risk Identification Framework for Foundation Model Uses cites this paper.

Developing a Risk Identification Framework for Foundation Model Uses Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:57.988743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:57.988743Z digest=sha256:6250e412fe090115c8439aaa6c60be37894287abe1b54efb797c558c39821135

Observation 9e5f17a6-a27d-47b9-b7eb-835e73519e26 · inbound

SCGAgent: Recreating the Benefits of Reasoning Models for Secure Code Generation with Agentic Workflows cites this paper.

SCGAgent: Recreating the Benefits of Reasoning Models for Secure Code Generation with Agentic Workflows Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:42:54.911928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:42:54.911928Z digest=sha256:cccece429e5b32a67501b12f018c4ea9853214a37d958256d5ec97af1cae5f96

Observation 78188daa-cfd8-411e-b556-0bf09ef0c4f5 · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.763796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.763796Z digest=sha256:bd2d60810ccf317f2e07878a04b815df34fa2a3cad29532abb5b9574bf09f69d

Observation 381c30dd-e45d-4868-af24-85cbb8ecd9f9 · inbound

Guiding AI to Fix Its Own Flaws: An Empirical Study on LLM-Driven Secure Code Generation cites this paper.

Guiding AI to Fix Its Own Flaws: An Empirical Study on LLM-Driven Secure Code Generation Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:56:13.509418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:56:13.509418Z digest=sha256:6235eb12aa0d933f5303c63a1ecfa82248ae891d1e5ee1686725fa40258899e4

Observation cf6837e0-a1cf-456c-b41d-df3c068a6907 · inbound

MGC: A Compiler Framework Exploiting Compositional Blindness in Aligned LLMs for Malware Generation cites this paper.

MGC: A Compiler Framework Exploiting Compositional Blindness in Aligned LLMs for Malware Generation Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:53.247909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:53.247909Z digest=sha256:4699e618ac6712cd08c47b1b6a3442c02aa7424bc403a6ef4870ef7265df94c8

Observation 97794215-ada6-49f4-9f3e-40a02bb0b665 · inbound

Understanding the Supply Chain and Risks of Large Language Model Applications cites this paper.

Understanding the Supply Chain and Risks of Large Language Model Applications Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:42.847281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:42.847281Z digest=sha256:788fbd8e77d83ab265d9884f6bece67a7b7cfbda183e266b13d3c89c2d7c9208

Observation 5dfe7ef9-3eb8-4fa5-8515-d47d7e71faec · inbound

Running in CIRCLE? A Simple Benchmark for LLM Code Interpreter Security cites this paper.

Running in CIRCLE? A Simple Benchmark for LLM Code Interpreter Security Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:25:15.757169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:25:15.757169Z digest=sha256:34669ef93ac5769a6c256e203fb23c06399eadb44b794642dd0fd62d5ee76fca

Observation 24a861b7-a4e2-4ef7-8c4e-c1211673c999 · inbound

Towards terahertz nanomechanics cites this paper.

Towards terahertz nanomechanics Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T01:03:49.483750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T01:03:49.483750Z digest=sha256:6aef0d16d0d4033c0d4e2e3eb7b2c89e00f758a6868bb43b2b879fa614505293

Observation 75bcb548-44e5-45d4-8578-254e6d951780 · inbound

ASTRA: Autonomous Spatial-Temporal Red-teaming for AI Software Assistants cites this paper.

ASTRA: Autonomous Spatial-Temporal Red-teaming for AI Software Assistants Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T01:04:32.435944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T01:04:32.435944Z digest=sha256:45e7dd297c20b2810eb7291e8a0c0160cfaa08f9314d2f2f16c34b36572ee66d

Observation 1f80b99b-ddea-4efa-bee6-62a473fe1d3f · inbound

Secure Code Generation at Scale with Reflexion cites this paper.

Secure Code Generation at Scale with Reflexion Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T23:50:49.056328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:50:49.056328Z digest=sha256:58f5b6745cadab884e7cb6e5556a0e03eb1cc259975bf72e6b4a576dcd374055

Observation 375b5ec1-aada-443f-8522-c7944673692a · inbound

BEAVER: An Efficient Deterministic LLM Verifier cites this paper.

BEAVER: An Efficient Deterministic LLM Verifier Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T01:58:51.237959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T01:58:44.719715Z digest=sha256:d9383ce296f8449e1a2c52a78936a6aefa8e4d038de8ec0003c68750a6569d85

Observation 12076545-9314-47ab-8e36-96d721f8ed1e · inbound

Extracting Recurring Vulnerabilities from Black-Box LLM-Generated Software cites this paper.

Extracting Recurring Vulnerabilities from Black-Box LLM-Generated Software Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T05:18:42.953661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:18:42.953661Z digest=sha256:3f967eff9ff7f96a268d5850a7bd9d7ff63dadf02b151f8c7316e05ddb6a1232

Observation 1c6d0de2-b199-4885-869e-e2348c008998 · inbound

"Tab, Tab, Bug": Security Pitfalls of Next Edit Suggestions in AI-Integrated IDEs cites this paper.

"Tab, Tab, Bug": Security Pitfalls of Next Edit Suggestions in AI-Integrated IDEs Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:57:29.297748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T06:53:13.235588Z digest=sha256:f973cc55a49edb61fd0e45f07f533810abd57be759e350d2f602b2634c4be424

Observation 94d15de5-8ac2-4d03-9fd8-6d3157135cbb · inbound

Broken by Default: A Formal Verification Study of Security Vulnerabilities in AI-Generated Code cites this paper.

Broken by Default: A Formal Verification Study of Security Vulnerabilities in AI-Generated Code Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:10:48.566036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T20:12:03.118074Z digest=sha256:05388e3e77da52c5e04f3f0ecf38e0d179408141dd383d9ded5375b1b62f9e16

Observation bebf90b3-ee10-42c9-9deb-daf54a51f640 · inbound

Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types cites this paper.

Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:31:00.338647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:08:25.471462Z digest=sha256:6349233af4994b6c1c9e141947ac276cac61cde3658b3afa18d5e97028dc37fc

Observation 7fbbb1e4-aa8e-49f4-b162-b551afa5aa3e · inbound

Adversarial Arena: Crowdsourcing Data Generation through Interactive Competition cites this paper.

Adversarial Arena: Crowdsourcing Data Generation through Interactive Competition Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:10:09.437088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T04:55:43.987116Z digest=sha256:ca72a4983aa4b9aad0656107543220e8cc34a85c7bb40289eb9a6112b6b1e8b9

Observation 2c5d9f52-8c89-4d6e-9660-35e553710ce1 · inbound

Towards Optimal Agentic Architectures for Offensive Security Tasks cites this paper.

Towards Optimal Agentic Architectures for Offensive Security Tasks Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:16:03.146927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T04:02:04.359269Z digest=sha256:5a58b1d81480ea8d08966018e73fe84a117a581fa6d72a926101eb851b792123

Observation 3eda2f9c-237c-4b36-803e-1bdc8981c0d5 · inbound

A Validated Prompt Bank for Malicious Code Generation: Separating Executable Weapons from Security Knowledge in 1,554 Consensus-Labeled Prompts cites this paper.

A Validated Prompt Bank for Malicious Code Generation: Separating Executable Weapons from Security Knowledge in 1,554 Consensus-Labeled Prompts Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:40:43.797249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T18:11:29.066362Z digest=sha256:b3bdee7caf0c300c62bfa634d1a6444ccf1e9406a9c4351c23fb907077c7ede7

Observation 89af28ed-bcf5-40d3-8611-5630b43ad2ef · inbound

Redefining AI Red Teaming in the Agentic Era: From Weeks to Hours cites this paper.

Redefining AI Red Teaming in the Agentic Era: From Weeks to Hours Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:51:44.776533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T16:06:18.057868Z digest=sha256:07214e1718fede16f6f34df96a4039a354fbdb0fe0014312a9a12c7468eb55ed

Observation 15bad8b2-2d34-44a3-928e-b4a531ea6e94 · inbound

Bridging Generation and Training: A Systematic Review of Quality Issues in LLMs for Code cites this paper.

Bridging Generation and Training: A Systematic Review of Quality Issues in LLMs for Code Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:26:04.477648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T17:37:51.790000Z digest=sha256:3c1edcc4f870afd51ce74d8b38f019c01299d80073a3447381e3ddf1ba4a11b5

Observation 593fda4e-e88b-4287-a655-07e4b936771f · inbound

SecureForge: Finding and Preventing Vulnerabilities in LLM-Generated Code via Prompt Optimization cites this paper.

SecureForge: Finding and Preventing Vulnerabilities in LLM-Generated Code via Prompt Optimization Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:26:24.588493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:09:08.378040Z digest=sha256:a0322ab03e748c9b2ed302d8234157bf834957e7ab3943fd813ba771025d8fc9

Observation d2708aee-a774-4683-95be-2aeb7caa45d7 · inbound

LLM-Agnostic Semantic Representation Attack cites this paper.

LLM-Agnostic Semantic Representation Attack Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:21:24.222423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:14:08.629862Z digest=sha256:ddac151f8e5aa90de04ecba1f53c2b9fff5d5f8ba925402f3f9d01aff01c1b29

Observation de0b055b-e787-4943-bf8c-d0ea299bb904 · inbound

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications cites this paper.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-19T23:32:52.541609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:8531980c983df94c71ff61d7e4f8130e946e51888fff7798044280878e1e6c6b

Observation e9453341-ea23-434c-9bd0-857ce0aca0b2 · inbound

Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025) cites this paper.

Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025) Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T07:39:48.536733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T07:38:43.363709Z digest=sha256:f0854bbb7a773a54612a54077962f7352688a3a064682e661010ceabad9c9728

Observation 014b4ae9-6471-42c6-9e79-be13ffc29e50 · inbound

HIDBench: Benchmarking Large Language Models for Host-Based Intrusion Detection cites this paper.

HIDBench: Benchmarking Large Language Models for Host-Based Intrusion Detection Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T08:54:45.868682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:52:34.079804Z digest=sha256:16ef156f680b3e71df6c35c7bd709f96351e1997f57c16abc82c6216e654b566

Observation 07a96601-9bbd-4e36-bdfd-14c9a7bae55b · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:07.706475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T05:50:28.114140Z digest=sha256:71722c08218fc6131ed4c4e1df5eb57e71714c7e453e1b7c5eb8516ba9a4522d

Observation 50142792-7fa7-41bc-bfa7-d3d3e189d660 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:06:43.062405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:05:27.736494Z digest=sha256:d456b5f17ce304d05f6eea1257b79be2b502e362d28bb201099c73739b7f9bd0

Observation 8497cf12-ab6e-4c82-9ac5-566da15afb6a · inbound

Security of LLM-generated Code: A Comparative Analysis cites this paper.

Security of LLM-generated Code: A Comparative Analysis Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:16:39.185029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:16:26.372764Z digest=sha256:e7828cb07514c7d383b2b9ca103112c8c732db3173d83916f74ff4d7d8283b6e

Observation 0e47e741-ab83-4a5a-aa3c-4c83b18a6e4d · inbound

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces cites this paper.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:36:15.066149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:3bb702286afcc9d97f735dccc7273d78456e72447332b874434a7f93088701de

Observation e7a4f50f-7229-44a5-a12a-25664d9e796a · inbound

Helpful or Harmful? Evaluating LLM-Assisted Vulnerability Patching via a Human Study cites this paper.

Helpful or Harmful? Evaluating LLM-Assisted Vulnerability Patching via a Human Study Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T21:10:09.288469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-25T19:02:45.109478Z digest=sha256:665cabc8d5e799d47c65d1f58273e6d83ceb3feb8815c404984b94a5e91329fd

Observation 9d048307-d19b-422f-a380-eb09b9fd25ee · inbound

Direct Causation in International Humanitarian Law and the Challenge of AI-Mediated Civilian Cyber Operations cites this paper.

Direct Causation in International Humanitarian Law and the Challenge of AI-Mediated Civilian Cyber Operations Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:54:22.125211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T07:49:28.819402Z digest=sha256:e495535bc0b1454fce0853074f7adf629c955fd9fe403d75e5242497505cce6d

Observation f2caab62-2233-4b1d-a14f-600eca1074a7 · inbound

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale cites this paper.

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T07:33:43.015966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:33:43.015966Z digest=sha256:ed366085032155a0ce06eae29fe9ae82ca4623ca19354989d8a27dba5e050075

Observation 0d6179b5-8f42-4710-af49-7b22b97a1710 · inbound

Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents cites this paper.

Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-11T22:40:37.839133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:40:37.839133Z digest=sha256:0f2044c74613e4e037634a94208f58695299712d79dece341c249ed27d893973

Observation c4b6b2a4-983a-46fb-a945-a3a224793c59 · inbound

Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents cites this paper.

Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T07:01:49.222325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T07:01:49.222325Z digest=sha256:1235cb3949560dbff0208d25432342820b0d06ec16091a33e4f882f10e430ec5

Observation 237e36ec-cebd-45cf-9c43-8a1d2b584881 · inbound

Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response cites this paper.

Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T02:40:29.099157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:40:29.099157Z digest=sha256:ea1c58267f4469ec0b19bf19aa99e777ab9bc52313f65ec0f7e5d1ebcb808037

Observation c2eb78e9-4e54-4e10-a6a0-07da0442523b · inbound

Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response cites this paper.

Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T01:33:17.643282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:33:17.643282Z digest=sha256:08adacb5fac2b39e450b61838cc9ae2075d8e5725bfeb6a240b7f7ab6aae0906

Observation 7cdc2bd7-d8d6-4bb1-8996-da6ea68649a7 · inbound

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) cites this paper.

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:17.193944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:17.193944Z digest=sha256:432203cac43574edea26fe90a4558541557bd9f49936f1bffa2cdc5576f23a14