Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T16:38:42.962606Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2411.13323.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T16:38:42.962606Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:20:21.052795Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-11T00:16:15.729268Z
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 19cccd37-eae3-4fe3-9883-69a6b40c60dc · outbound
Are Large Language Models Memorizing Bug Benchmarks? A survey on software fault localization,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 619004af-2b05-49c6-aad8-7356b5ef0525 · outbound
Are Large Language Models Memorizing Bug Benchmarks? Automated Program Repair,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9546f144-40e9-40a9-be7a-788c08e8949b · outbound
Are Large Language Models Memorizing Bug Benchmarks? Defects4j: a database of existing faults to enable controlled testing studies for java programs,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a52e983f-e0fe-4a17-b62e-0a3a1e26410c · outbound
Are Large Language Models Memorizing Bug Benchmarks? Bugsinpy: a database of existing bugs in python programs to enable controlled testing and debugging studies,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d0bfa93d-fe52-4f03-8c37-1e0cf92ad63f · outbound
Are Large Language Models Memorizing Bug Benchmarks? SWE-bench: Can language models resolve real-world github issues?
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 05131f45-ffc5-4890-a61f-3b941cf05c6b · outbound
Are Large Language Models Memorizing Bug Benchmarks? Concerned with Data Contamination? Assessing Countermeasures in Code Language Model
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1c4d370-6e1e-4cf5-bcab-de488e646eed · outbound
Are Large Language Models Memorizing Bug Benchmarks? Leakage and the Reproducibility Crisis in ML-based Science
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91903aa1-6de5-4571-be83-ce728ec2225b · outbound
Are Large Language Models Memorizing Bug Benchmarks? Benchmarking Benchmark Leakage in Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f9e806b-50c5-4845-a892-601a9d8ad91c · outbound
Are Large Language Models Memorizing Bug Benchmarks? Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cdb5a41-7298-4f92-9321-46d76c4fce2f · outbound
Are Large Language Models Memorizing Bug Benchmarks? Bugsc++: A highly usable real world defect benchmark for c/c++,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d09c4910-7504-420a-9664-9fc61b9e9c9c · outbound
Are Large Language Models Memorizing Bug Benchmarks? Gitbug-java: A reproducible benchmark of recent java bugs,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7fb3822f-891b-4631-97b0-6ce0bed8fcee · outbound
Are Large Language Models Memorizing Bug Benchmarks? Agentless: Demystifying LLM-based Software Engineering Agents
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9db61581-7331-4858-9b14-7e0d47a8eb4c · outbound
Are Large Language Models Memorizing Bug Benchmarks? OpenHands: An Open Platform for AI Software Developers as Generalist Agents
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 734c1bc4-4623-40cb-aa05-49b7d4d376e7 · outbound
Are Large Language Models Memorizing Bug Benchmarks? SWE-Bench+: Enhanced Coding Benchmark for LLMs
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6e344e4-9bad-4ae1-9a67-b4b9258ae866 · outbound
Are Large Language Models Memorizing Bug Benchmarks? On the resemblance and containment of documents,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6242c7ba-190b-4661-ac2c-3fde315d8ec7 · outbound
Are Large Language Models Memorizing Bug Benchmarks? Similarity search in high dimensions via hashing,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5233ba2b-6590-4991-bb27-c1ab79d191d5 · outbound
Are Large Language Models Memorizing Bug Benchmarks? Codegen: An open large language model for code with multi-turn program synthesis,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fcf00b6b-9882-4094-a865-574eb4dff3d6 · outbound
Are Large Language Models Memorizing Bug Benchmarks? Code llama: Open foundation models for code,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5722db26-a0a7-4e0f-b402-92c10b750856 · outbound
Are Large Language Models Memorizing Bug Benchmarks? Llama: Open and efficient foundation language models,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b7abd74-868e-4562-a69b-ed465abbc3a7 · outbound
Are Large Language Models Memorizing Bug Benchmarks? Code Llama: Open Foundation Models for Code
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5648e50f-7837-4773-8691-f79a815bf547 · outbound
Are Large Language Models Memorizing Bug Benchmarks? Gemma 2: Improving Open Language Models at a Practical Size
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a8e0c35-14a6-47d8-9314-ecbfd02df6bf · outbound
Are Large Language Models Memorizing Bug Benchmarks? Starcoder 2 and the stack v2: The next generation,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c97a4665-f0e4-47f3-a172-c052093eb1fa · outbound
Are Large Language Models Memorizing Bug Benchmarks? Mistral 7B
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29f57848-3c77-446d-ac1a-2e9bf18ee480 · outbound
Are Large Language Models Memorizing Bug Benchmarks? Codegemma: Open code models based on gemma,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5f6b8255-4788-4973-bdc5-e68bb0434e98 · outbound
Are Large Language Models Memorizing Bug Benchmarks? Repairllama: Efficient representations and fine-tuned adapters for program repair,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1a7cc93e-7ad8-43a2-8bcb-3823f9b331c6 · outbound
Are Large Language Models Memorizing Bug Benchmarks? Large language model for vulnerability detection: Emerging results and future directions,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 10514c65-8456-4621-9750-548c46acfb89 · outbound
Are Large Language Models Memorizing Bug Benchmarks? Large language models for test-free fault localization,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 51d34f47-928c-4ef0-b6c9-2fea25c1c71e · outbound
Are Large Language Models Memorizing Bug Benchmarks? A study of the uniqueness of source code,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 59e1695b-2a69-4bf8-ab28-b7dafbd097c9 · outbound
Are Large Language Models Memorizing Bug Benchmarks? RepairLLaMA: Efficient Representations and Fine-Tuned Adapters for Program Repair
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49ac9234-caa2-4cd3-aa83-76a8f17f36c4 · outbound
Are Large Language Models Memorizing Bug Benchmarks? Memorization without overfitting: Analyzing the training dynamics of large language models,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 41c367a4-fe6c-4c5f-af46-ae0e0176eb1d · outbound
Are Large Language Models Memorizing Bug Benchmarks? The stack: 3 TB of permissively licensed source code,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5b037e62-a0b0-4319-9405-f770df87ecad · outbound
Are Large Language Models Memorizing Bug Benchmarks? Measuring massive multitask language understanding,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation aaaadc1c-7b6a-4bde-b9b6-f7a4fdaaff5e · outbound
Are Large Language Models Memorizing Bug Benchmarks? An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccc78991-03b9-4c7d-9762-de97282f6fd7 · outbound
Are Large Language Models Memorizing Bug Benchmarks? Program Synthesis with Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1acba78-afc4-4fa8-8b08-09bdf13a6c33 · outbound
Are Large Language Models Memorizing Bug Benchmarks? Squad: 100, 000+ questions for machine comprehension of text,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d5cf801-2987-42b0-9aff-464a161cbd97 · outbound
Are Large Language Models Memorizing Bug Benchmarks? Keep the Conversation Going: Fixing 162 out of 337 bugs for $0.42 each using ChatGPT
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3a57390-4812-4472-819f-e5750cd47e0a · outbound
Are Large Language Models Memorizing Bug Benchmarks? Evaluating large language models trained on code,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b293650d-04a7-415b-9989-f12de4f16fe7 · outbound
Are Large Language Models Memorizing Bug Benchmarks? LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61797f08-cfb1-4f59-b0fd-409ea0d30555 · outbound
Are Large Language Models Memorizing Bug Benchmarks? Why does your data leak? uncovering the data leakage in cloud from mobile apps,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 47de63b2-7354-4d2d-944e-93e212e32bf4 · outbound
Are Large Language Models Memorizing Bug Benchmarks? Measuring coding challenge competence with APPS,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation aa29b1cc-1416-4b5b-b1ef-8677991695f3 · outbound
Are Large Language Models Memorizing Bug Benchmarks? Detecting pretraining data from large language models,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3b338049-eaca-4c1a-bed3-2f8252848170 · outbound
Are Large Language Models Memorizing Bug Benchmarks? LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 371f832b-16fd-45fa-8db1-91e2625b18a2 · outbound
Are Large Language Models Memorizing Bug Benchmarks? Evaluating Large Language Models Trained on Code
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 860f927a-9d1b-4c35-b4fc-9d5cea66c2cc · outbound
Are Large Language Models Memorizing Bug Benchmarks? Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df50b642-a106-4b59-b337-7ca5e5baebda · outbound
Are Large Language Models Memorizing Bug Benchmarks? CodeGemma: Open Code Models Based on Gemma
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 904278f3-d300-46fa-94f9-2b1ea3fc9164 · inbound
CveBinarySheet: A Comprehensive Pre-built Binaries Database for IoT Vulnerability Analysis Are Large Language Models Memorizing Bug Benchmarks?
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.