Pith. sign in

Paper Citation Record · LEDGER

Are Large Language Models Memorizing Bug Benchmarks?

As of 13 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2411.13323.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13323 v3

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:38:42.962606Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:20:21.052795Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-11T00:16:15.729268Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy22
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 19cccd37-eae3-4fe3-9883-69a6b40c60dc · outbound

This paper cites A survey on software fault localization,.

Are Large Language Models Memorizing Bug Benchmarks? A survey on software fault localization,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:44.177178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:38:42.252494Z digest=sha256:2cac6a072d39a6d5f95219f3e2ed345e393efa8baeabab0960fc7384f6c72ff6

Observation 619004af-2b05-49c6-aad8-7356b5ef0525 · outbound

This paper cites Automated Program Repair,.

Are Large Language Models Memorizing Bug Benchmarks? Automated Program Repair,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:44.163595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:38:42.256974Z digest=sha256:6d8227a62f27265ea565c19b291117fddd2c2e4c49484349501b541aca5b7a6d

Observation 9546f144-40e9-40a9-be7a-788c08e8949b · outbound

This paper cites Defects4j: a database of existing faults to enable controlled testing studies for java programs,.

Are Large Language Models Memorizing Bug Benchmarks? Defects4j: a database of existing faults to enable controlled testing studies for java programs,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:44.151216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:38:42.261686Z digest=sha256:64202b89c2e66a189857fce5458efc4b5dfc64fcc5b506a4f13d1bcc81ba1f1d

Observation a52e983f-e0fe-4a17-b62e-0a3a1e26410c · outbound

This paper cites Bugsinpy: a database of existing bugs in python programs to enable controlled testing and debugging studies,.

Are Large Language Models Memorizing Bug Benchmarks? Bugsinpy: a database of existing bugs in python programs to enable controlled testing and debugging studies,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:44.139126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:38:42.265987Z digest=sha256:63e5c796822516df311e68e3e75ca47a132a88bee3b78f78bf1ade71901334fa

Observation d0bfa93d-fe52-4f03-8c37-1e0cf92ad63f · outbound

This paper cites SWE-bench: Can language models resolve real-world github issues?.

Are Large Language Models Memorizing Bug Benchmarks? SWE-bench: Can language models resolve real-world github issues?

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:44.124917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:38:42.270348Z digest=sha256:cfbb00af5f766e1d689509300708ee08e3bec422a64a7f432a0b2f29ede2b116

Observation 05131f45-ffc5-4890-a61f-3b941cf05c6b · outbound

This paper cites Concerned with Data Contamination? Assessing Countermeasures in Code Language Model.

Are Large Language Models Memorizing Bug Benchmarks? Concerned with Data Contamination? Assessing Countermeasures in Code Language Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.274919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.274919Z digest=sha256:2a08971ff59aaa5d560ea8807c4f8ac8798424fa4a01760eb16ac9999b257038

Observation e1c4d370-6e1e-4cf5-bcab-de488e646eed · outbound

This paper cites Leakage and the Reproducibility Crisis in ML-based Science.

Are Large Language Models Memorizing Bug Benchmarks? Leakage and the Reproducibility Crisis in ML-based Science

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.280170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.280170Z digest=sha256:3cdcd531016611a5156160496ca17da91cdd9cbd55071c776392fd7a8d9dd43e

Observation 91903aa1-6de5-4571-be83-ce728ec2225b · outbound

This paper cites Benchmarking Benchmark Leakage in Large Language Models.

Are Large Language Models Memorizing Bug Benchmarks? Benchmarking Benchmark Leakage in Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.284289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.284289Z digest=sha256:2c18e0076afac830566e84895ebbebc180a537be21a3d2dd9e5ea89bb6f0083b

Observation 4f9e806b-50c5-4845-a892-601a9d8ad91c · outbound

This paper cites Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation.

Are Large Language Models Memorizing Bug Benchmarks? Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.288484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.288484Z digest=sha256:ea69d0b3cff08d2ee768a22baa3337394c5d004845453ce3a1366b693d3ebf88

Observation 9cdb5a41-7298-4f92-9321-46d76c4fce2f · outbound

This paper cites Bugsc++: A highly usable real world defect benchmark for c/c++,.

Are Large Language Models Memorizing Bug Benchmarks? Bugsc++: A highly usable real world defect benchmark for c/c++,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.296648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.296648Z digest=sha256:1bd97cddac3489e01645da0557a0e15aa56d1c7ecde53aeac2ae383438623117

Observation d09c4910-7504-420a-9664-9fc61b9e9c9c · outbound

This paper cites Gitbug-java: A reproducible benchmark of recent java bugs,.

Are Large Language Models Memorizing Bug Benchmarks? Gitbug-java: A reproducible benchmark of recent java bugs,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.959940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:38:42.300069Z digest=sha256:1bb06f00cd2eb30619a973a9a8a28fd8020d21b23454fceae0e6fa69a68b1cfb

Observation 7fb3822f-891b-4631-97b0-6ce0bed8fcee · outbound

This paper cites Agentless: Demystifying LLM-based Software Engineering Agents.

Are Large Language Models Memorizing Bug Benchmarks? Agentless: Demystifying LLM-based Software Engineering Agents

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.304045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.304045Z digest=sha256:9bfab447c2d89f074fe161eace5c9df057cb39207bec2e165657a94534fa82cd

Observation 9db61581-7331-4858-9b14-7e0d47a8eb4c · outbound

This paper cites OpenHands: An Open Platform for AI Software Developers as Generalist Agents.

Are Large Language Models Memorizing Bug Benchmarks? OpenHands: An Open Platform for AI Software Developers as Generalist Agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.308840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.308840Z digest=sha256:29512f7cac3eb8b49df8ee7ff3e87ab207207b859ed44a33e8168e9063e2a78a

Observation 734c1bc4-4623-40cb-aa05-49b7d4d376e7 · outbound

This paper cites SWE-Bench+: Enhanced Coding Benchmark for LLMs.

Are Large Language Models Memorizing Bug Benchmarks? SWE-Bench+: Enhanced Coding Benchmark for LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.349704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.349704Z digest=sha256:f117e7ade70450b7921dabfb6153b9810a6ee42ebc5abea1208dc6249d2dc3c5

Observation c6e344e4-9bad-4ae1-9a67-b4b9258ae866 · outbound

This paper cites On the resemblance and containment of documents,.

Are Large Language Models Memorizing Bug Benchmarks? On the resemblance and containment of documents,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.877092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:38:42.405481Z digest=sha256:757794116457550bad438a3369c240e49ea3843d2c890ab9d6d912f14bf5a8a8

Observation 6242c7ba-190b-4661-ac2c-3fde315d8ec7 · outbound

This paper cites Similarity search in high dimensions via hashing,.

Are Large Language Models Memorizing Bug Benchmarks? Similarity search in high dimensions via hashing,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.863930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:38:42.506527Z digest=sha256:49d32a668910dea2f113beba657fb85026c1ff9ce6d0716f76e1b9e456d2f2b2

Observation 5233ba2b-6590-4991-bb27-c1ab79d191d5 · outbound

This paper cites Codegen: An open large language model for code with multi-turn program synthesis,.

Are Large Language Models Memorizing Bug Benchmarks? Codegen: An open large language model for code with multi-turn program synthesis,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.852055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:38:42.510429Z digest=sha256:d7a91574854e89f8f2c79f382ff66a64213a622aa091cde6692ad4d581932182

Observation fcf00b6b-9882-4094-a865-574eb4dff3d6 · outbound

This paper cites Code llama: Open foundation models for code,.

Are Large Language Models Memorizing Bug Benchmarks? Code llama: Open foundation models for code,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.837616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:38:42.514409Z digest=sha256:49d12dfd964461e071dc0b402690ea289b7dac946700e8b7c5f553882c5de159

Observation 5722db26-a0a7-4e0f-b402-92c10b750856 · outbound

This paper cites Llama: Open and efficient foundation language models,.

Are Large Language Models Memorizing Bug Benchmarks? Llama: Open and efficient foundation language models,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.522310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.522310Z digest=sha256:9871c331a561ee74bc6faab9b684426529c00bcd41cec51c662b3c68add85494

Observation 6b7abd74-868e-4562-a69b-ed465abbc3a7 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Are Large Language Models Memorizing Bug Benchmarks? Code Llama: Open Foundation Models for Code

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.518109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.518109Z digest=sha256:c718a87016ebb8be7e54d8ef60988ee79d2d6715c6d2d7b6d23933fc1c61459f

Observation 5648e50f-7837-4773-8691-f79a815bf547 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Are Large Language Models Memorizing Bug Benchmarks? Gemma 2: Improving Open Language Models at a Practical Size

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.530024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.530024Z digest=sha256:1affa032bd0c25ebe4368f43b23801d4e15f9705b36982e70272a09471357ba6

Observation 6a8e0c35-14a6-47d8-9314-ecbfd02df6bf · outbound

This paper cites Starcoder 2 and the stack v2: The next generation,.

Are Large Language Models Memorizing Bug Benchmarks? Starcoder 2 and the stack v2: The next generation,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.776154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:38:42.525883Z digest=sha256:3b94f73fe5337d7d3e084378577da8ab9c8d9586946d934695d05f14522baa9c

Observation c97a4665-f0e4-47f3-a172-c052093eb1fa · outbound

This paper cites Mistral 7B.

Are Large Language Models Memorizing Bug Benchmarks? Mistral 7B

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.758008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.758008Z digest=sha256:616119e734ef246ab8ca92670b71eafef9872e4e9f164571bedfe201ba9e2075

Observation 29f57848-3c77-446d-ac1a-2e9bf18ee480 · outbound

This paper cites Codegemma: Open code models based on gemma,.

Are Large Language Models Memorizing Bug Benchmarks? Codegemma: Open code models based on gemma,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.686661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:38:42.586465Z digest=sha256:316b930d70e0ae0cde1b4beb08524ec01eabcdfb91985620fa1557fa87a7a0e6

Observation 5f6b8255-4788-4973-bdc5-e68bb0434e98 · outbound

This paper cites Repairllama: Efficient representations and fine-tuned adapters for program repair,.

Are Large Language Models Memorizing Bug Benchmarks? Repairllama: Efficient representations and fine-tuned adapters for program repair,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.660234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:38:42.795695Z digest=sha256:792d484cf38083167580fd56c7f3fbdfe51e17795ec9515dfeb0b9412adb4a55

Observation 1a7cc93e-7ad8-43a2-8bcb-3823f9b331c6 · outbound

This paper cites Large language model for vulnerability detection: Emerging results and future directions,.

Are Large Language Models Memorizing Bug Benchmarks? Large language model for vulnerability detection: Emerging results and future directions,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.647047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:38:42.804097Z digest=sha256:e2d330a114035dd34b11f3cc3bbce1eec759e8fbda6d36ff86a59c4d48969e8a

Observation 10514c65-8456-4621-9750-548c46acfb89 · outbound

This paper cites Large language models for test-free fault localization,.

Are Large Language Models Memorizing Bug Benchmarks? Large language models for test-free fault localization,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.673850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:38:42.791692Z digest=sha256:47bff501236368a54337dc377fcd2df1bc34e8bc2a2ec23b8c2b6d517cb3c753

Observation 51d34f47-928c-4ef0-b6c9-2fea25c1c71e · outbound

This paper cites A study of the uniqueness of source code,.

Are Large Language Models Memorizing Bug Benchmarks? A study of the uniqueness of source code,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.468959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:38:42.811382Z digest=sha256:cae14970afd606d277867bdd136ed2206023faf6f39914f9c34025f1bb5ae05d

Observation 59e1695b-2a69-4bf8-ab28-b7dafbd097c9 · outbound

This paper cites RepairLLaMA: Efficient Representations and Fine-Tuned Adapters for Program Repair.

Are Large Language Models Memorizing Bug Benchmarks? RepairLLaMA: Efficient Representations and Fine-Tuned Adapters for Program Repair

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.799148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.799148Z digest=sha256:f4bf1ee96dcfcef333d855d9441e74d93940c36e2bbfdf609354aa442e558178

Observation 49ac9234-caa2-4cd3-aa83-76a8f17f36c4 · outbound

This paper cites Memorization without overfitting: Analyzing the training dynamics of large language models,.

Are Large Language Models Memorizing Bug Benchmarks? Memorization without overfitting: Analyzing the training dynamics of large language models,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.455722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:38:42.819025Z digest=sha256:400571bbc21d9e027a8ca4125ba60fae9bfebfab87b6c6289fae3207f1998c8c

Observation 41c367a4-fe6c-4c5f-af46-ae0e0176eb1d · outbound

This paper cites The stack: 3 TB of permissively licensed source code,.

Are Large Language Models Memorizing Bug Benchmarks? The stack: 3 TB of permissively licensed source code,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.575402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:38:42.807962Z digest=sha256:d5aedd2efc8ea0b1f72c12ed344055c437a347299408ab266bad6ea46f112a54

Observation 5b037e62-a0b0-4319-9405-f770df87ecad · outbound

This paper cites Measuring massive multitask language understanding,.

Are Large Language Models Memorizing Bug Benchmarks? Measuring massive multitask language understanding,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.440137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:38:42.826381Z digest=sha256:7cb10440edd71b4919b3817821df6d00b5a55d37e1e7bf451c77fbae6fc0efe0

Observation aaaadc1c-7b6a-4bde-b9b6-f7a4fdaaff5e · outbound

This paper cites An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning.

Are Large Language Models Memorizing Bug Benchmarks? An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.814911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.814911Z digest=sha256:ae82fc843bb1fd52fbba16aa318f25ccfc6f126bc94290c601755f6aa7ee6f04

Observation ccc78991-03b9-4c7d-9762-de97282f6fd7 · outbound

This paper cites Program Synthesis with Large Language Models.

Are Large Language Models Memorizing Bug Benchmarks? Program Synthesis with Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.838268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.838268Z digest=sha256:aca8d3bc792b71df8d992ea14740a1a08c1fabf9471145b0e950256099814fe3

Observation a1acba78-afc4-4fa8-8b08-09bdf13a6c33 · outbound

This paper cites Squad: 100, 000+ questions for machine comprehension of text,.

Are Large Language Models Memorizing Bug Benchmarks? Squad: 100, 000+ questions for machine comprehension of text,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.822872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.822872Z digest=sha256:107e42e95677ee332410b4ad491f2944997a7a407c26dc4948c56d17a6b17112

Observation 7d5cf801-2987-42b0-9aff-464a161cbd97 · outbound

This paper cites Keep the Conversation Going: Fixing 162 out of 337 bugs for $0.42 each using ChatGPT.

Are Large Language Models Memorizing Bug Benchmarks? Keep the Conversation Going: Fixing 162 out of 337 bugs for $0.42 each using ChatGPT

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.845327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.845327Z digest=sha256:3075fc3a4522a0decce33a833e5d27632e9a7a3874ad719e60c6cebe4943a439

Observation a3a57390-4812-4472-819f-e5750cd47e0a · outbound

This paper cites Evaluating large language models trained on code,.

Are Large Language Models Memorizing Bug Benchmarks? Evaluating large language models trained on code,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.830570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.830570Z digest=sha256:4169279828943ebefddeca0d93e5055a4a6ca3f87a95ba092b371393cfffa92b

Observation b293650d-04a7-415b-9989-f12de4f16fe7 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Are Large Language Models Memorizing Bug Benchmarks? LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.893040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.893040Z digest=sha256:587deda8395c67fdcee6ad9eab778605481bc866531e8c4bf49ecb1e6c5265f7

Observation 61797f08-cfb1-4f59-b0fd-409ea0d30555 · outbound

This paper cites Why does your data leak? uncovering the data leakage in cloud from mobile apps,.

Are Large Language Models Memorizing Bug Benchmarks? Why does your data leak? uncovering the data leakage in cloud from mobile apps,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.353535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:38:42.962606Z digest=sha256:552dfe466f087dac78ce85a9253f23d476b6e82a08d5fa03180fc24f1e964d94

Observation 47de63b2-7354-4d2d-944e-93e212e32bf4 · outbound

This paper cites Measuring coding challenge competence with APPS,.

Are Large Language Models Memorizing Bug Benchmarks? Measuring coding challenge competence with APPS,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.418581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:38:42.841580Z digest=sha256:ef52dd163ec1286917a1490d79ca0bef9972616aaef3825297ca3f693767727d

Observation aa29b1cc-1416-4b5b-b1ef-8677991695f3 · outbound

This paper cites Detecting pretraining data from large language models,.

Are Large Language Models Memorizing Bug Benchmarks? Detecting pretraining data from large language models,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:43.405143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:38:42.849344Z digest=sha256:d05479c1a49a58117e79031a24d6830b43e64d5fce8589b2c2c2b00b1584a687

Observation 3b338049-eaca-4c1a-bed3-2f8252848170 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Are Large Language Models Memorizing Bug Benchmarks? LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.958319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.958319Z digest=sha256:b54513bb75385bc72eae51ce678140a070bbe94d82f51c8090127c45dd1531f0

Observation 371f832b-16fd-45fa-8db1-91e2625b18a2 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Are Large Language Models Memorizing Bug Benchmarks? Evaluating Large Language Models Trained on Code

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.833805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.833805Z digest=sha256:645e48aeeed1c4db87ecc59cd898423f9098baa1bf9d7e98739bfa6b75eda6cd

Observation 860f927a-9d1b-4c35-b4fc-9d5cea66c2cc · outbound

This paper cites Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation.

Are Large Language Models Memorizing Bug Benchmarks? Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.292215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.292215Z digest=sha256:42f80b74f0f18a45e6f6af8b573b0279db5e4991d0cdd2b183af23cd22cea258

Observation df50b642-a106-4b59-b337-7ca5e5baebda · outbound

This paper cites CodeGemma: Open Code Models Based on Gemma.

Are Large Language Models Memorizing Bug Benchmarks? CodeGemma: Open Code Models Based on Gemma

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.656619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.656619Z digest=sha256:a3a72551580ef21f33567b09d9623fd6a85921b4ab0c93720cd97b7f331ed6e2

Pith citing papers

Observation 904278f3-d300-46fa-94f9-2b1ea3fc9164 · inbound

CveBinarySheet: A Comprehensive Pre-built Binaries Database for IoT Vulnerability Analysis cites this paper.

CveBinarySheet: A Comprehensive Pre-built Binaries Database for IoT Vulnerability Analysis Are Large Language Models Memorizing Bug Benchmarks?

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:20:21.116359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:20:21.052795Z digest=sha256:6e62fbe646cf46a01f886b256c07d8f20366e95de6b7c4390d1322947a553c80