Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T19:04:26.024098Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 11 inbound Pith citation observations for arXiv:2501.10711.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T19:04:26.024098Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:25:46.325101Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T08:17:45.313006Z
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 34896d21-bd5d-4b53-8da1-2aee4c39f741 · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility Program Synthesis with Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 307726f7-db38-4df1-aa70-44029ce64160 · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7fb5baa-a027-43a4-b79a-749c7f0ae816 · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility On the Impacts of Contexts on Repository-Level Code Generation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ca4f069-33a3-4768-8b75-1b2ff9c01810 · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility Execution-based Evaluation for Data Science Code Generation Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 392285b3-6eaf-4498-b896-a7b3e5ae6a71 · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility RepoQA: Evaluating Long Context Code Understanding
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da9dd4fe-2709-43c9-adad-ac262a9c5c87 · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility The RealHumanEval: Evaluating Large Language Models' Abilities to Support Programmers
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0985be0-94b7-4b92-a3e6-5fa8d74508c8 · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility RunBugRun -- An Executable Dataset for Automated Program Repair
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4957f80-38f9-426e-86cb-37a07e3b7a1a · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 841a835e-fbcf-427d-a7d2-e0249605e945 · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility ACM Comput
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e25778f0-993d-4db0-91f9-63d64667c6ec · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5538b631-824e-4882-be67-1604f0ec4d8a · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility AMBROSIA: A Benchmark for Parsing Ambiguous Questions into Database Queries
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 641b435f-c8d6-4d26-aeaf-5bf8f6da3200 · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41fb8a0d-318b-45b7-95ab-53e4389c1f5e · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility Benchmarking TPU, GPU, and CPU Platforms for Deep Learning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c0761f5-e342-418f-a39e-575223e02759 · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8871a011-9d4f-4647-b182-379e83234441 · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility Plot2Code: A Comprehensive Benchmark for Evaluating Multi-modal Large Language Models in Code Generation from Scientific Plots
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5a330f3-866f-42e8-8471-0b3810afb14f · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility CRUXEval-X: A Benchmark for Multilingual Code Reasoning, Understanding and Execution
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c618f911-c028-42e7-98f4-592bb192423a · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility HumanEvo: An Evolution-aware Benchmark for More Realistic Evaluation of Repository-level Code Generation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bbc60dd-80a8-42d0-82c1-d7fec200b412 · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility 0 1000 2000 3000 0 10 20 30 40 50 60 70 80 90 100 110 120 130 140 150 160 170 180 190 Citations Figure 10: Citation Distribution of Benchmarks Coding Task
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7144b574-bcae-4541-bb62-9410db5821aa · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility For example, Hu- manEval (Chen et al., 2021a) and MBPP (Austin et al., 2021)), class-level (i.e., a class with mul- tiple function units of code
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e886fbfb-412f-468a-8bd7-cdf61fa6333e · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility "" if len(dict.keys()) == 0: return False else: state =
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e2308edd-b6d1-4238-a348-5085acc9db28 · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility assert upper_ctr('PYthon') == 1
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a0345340-8e1f-40de-b16c-c684c63cdb84 · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis
Reference 294
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b87d9451-8b50-4d17-aa66-36f687992189 · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility RES-Q: Evaluating Code-Editing Large Language Model Systems at the Repository Scale
Reference 516
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6e9b65b5-40ce-4d6e-bbf6-3c801369c065 · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility The Impact of Reasoning Step Length on Large Language Models
Reference 1442
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24b41b94-ab75-4ba6-a038-b2b2c6862b19 · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility Gorilla: Large Language Model Connected with Massive APIs
Reference 1569
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9220f3cd-78d5-4726-b8e8-7a63997dcb24 · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility VHDL-Eval: A Framework for Evaluating Large Language Models in VHDL Code Generation
Reference 2014
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8db7c4e9-ea1e-4336-94e0-7ecb703eb50f · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7457eae2-0314-4877-87a8-8e5fa287585b · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility Dianshu Liao, Shidong Pan, Xiaoyu Sun, Xiaoxue Ren, Qing Huang, Zhenchang Xing, Huan Jin, and Qinying Li
Reference 2018
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c77e4a68-9820-48cf-837b-9db14ee61242 · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility Unresolved cited work
Reference 2019
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bc2a4932-c16a-47b3-844d-063958e4e60f · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility In Proceedings of the 28th International Confer- ence on Computational Linguistics, COLING 2020, Barcelona, Spain (Online), December 8-13, 2020 , pages 26–38
Reference 2020
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5439fe26-5714-40ce-9e05-b0c5bb2310f4 · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility measure the ability of these models to synthesize short Python programs from natural language descriptions
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9e39f514-2992-4811-9d48-abf6cf58bbcd · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility SySeVR: A Framework for Using Deep Learning to Detect Software Vulnerabilities
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cc6295b4-01b7-4629-b9f5-3c910c840f3c · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility In 45th IEEE/ACM International Conference on Software Engineering, ICSE 2023, Melbourne, Australia, May 14-20, 2023, pages 1430–
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b154828c-19f1-4b79-84bf-0bb7265d9243 · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility Long Code Arena: a Set of Benchmarks for Long-Context Code Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43d11c34-c450-4efa-a791-c2a118d34726 · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility Lakshya A
Reference 5445
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 35c6a7bb-d39d-49b3-a7d7-83b04d35f4e5 · outbound
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility Ramakrishna Bairi, Atharv Sonwane, Aditya Kanade, Vageesh D
Reference 8474
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation beede259-c224-4d6f-b773-850b7dbae6e6 · inbound
Across Programming Language Silos: A Study on Cross-Lingual Retrieval-augmented Code Generation Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3973a4fa-89d9-446d-a73e-f298d673980b · inbound
A Conceptual Framework for AI Capability Evaluations Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58917dc5-a8db-4654-8d3a-2a65a00bfc8e · inbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a12dde6-efe5-4dfd-a678-8026e8bc1c33 · inbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fda34f98-829b-403e-be0e-948799f4c56c · inbound
Guidelines for Empirical Studies in Software Engineering involving Large Language Models Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 65779e12-7cd7-49cc-bd71-1f333fd5523c · inbound
Guidelines for Empirical Studies in Software Engineering involving Large Language Models Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6631b1be-bbee-457c-9245-9f7bb63248dd · inbound
Knowledge-Graph-Driven Data Synthesis for Low-Resource Software Development: A HarmonyOS Case Study Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0061e5a0-8657-4136-9935-52d0364e35f4 · inbound
Edit, But Verify: An Empirical Audit of Instructed Code-Editing Benchmarks Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8513986c-d294-41b6-a051-13b72123b60d · inbound
Flaws in the LLM Automation Narrative Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cc18cf3f-27f6-45f0-a7e8-ac73c421df8c · inbound
AlgoBench: Benchmarking Algorithmic Adaptation in Code Generation Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation eb8b0150-4c21-4608-81bf-219d2a213365 · inbound
To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.