Pith. sign in

Paper Citation Record · LEDGER

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility

As of 12 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 11 inbound Pith citation observations for arXiv:2501.10711.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.10711 v5

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T19:04:26.024098Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:25:46.325101Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T08:17:45.313006Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 34896d21-bd5d-4b53-8da1-2aee4c39f741 · outbound

This paper cites Program Synthesis with Large Language Models.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility Program Synthesis with Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T19:04:25.901736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:04:25.901736Z digest=sha256:e2c86488b7e44bc0f78341f31a2f8479c3c43adc7c92f6519a140b6239e341ba

Observation 307726f7-db38-4df1-aa70-44029ce64160 · outbound

This paper cites CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T19:04:25.915749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:04:25.915749Z digest=sha256:e22370e243ba198a7ff24e05a702d7a939e3bd29b885749b57af9f84aacbac5a

Observation d7fb5baa-a027-43a4-b79a-749c7f0ae816 · outbound

This paper cites On the Impacts of Contexts on Repository-Level Code Generation.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility On the Impacts of Contexts on Repository-Level Code Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T19:04:25.919849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:04:25.919849Z digest=sha256:6827b9fe33b224373486ebd5ae6204d9ba5512e81f8f0873cd9417f3d57de545

Observation 8ca4f069-33a3-4768-8b75-1b2ff9c01810 · outbound

This paper cites Execution-based Evaluation for Data Science Code Generation Models.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility Execution-based Evaluation for Data Science Code Generation Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T19:04:25.923680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:04:25.923680Z digest=sha256:104c401ee83ab8e29f2d11646cb63794b5d2e23f0c8aa1efcb642ce5c09d0f35

Observation 392285b3-6eaf-4498-b896-a7b3e5ae6a71 · outbound

This paper cites RepoQA: Evaluating Long Context Code Understanding.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility RepoQA: Evaluating Long Context Code Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T19:04:25.946456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:04:25.946456Z digest=sha256:3768d6021e53f6d6b03746520d0db16c209a19377df617ebffc653825f1c435c

Observation da9dd4fe-2709-43c9-adad-ac262a9c5c87 · outbound

This paper cites The RealHumanEval: Evaluating Large Language Models' Abilities to Support Programmers.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility The RealHumanEval: Evaluating Large Language Models' Abilities to Support Programmers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T19:04:25.950216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:04:25.950216Z digest=sha256:d73e7ff219d6186acec9a7a8b343d8a6902183690907351e43b2c80893b3d0b5

Observation b0985be0-94b7-4b92-a3e6-5fa8d74508c8 · outbound

This paper cites RunBugRun -- An Executable Dataset for Automated Program Repair.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility RunBugRun -- An Executable Dataset for Automated Program Repair

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T19:04:25.961358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:04:25.961358Z digest=sha256:0e16f2755fb70f5a12053fdece3072482526198ad8dca167ad5df40e79e0f827

Observation f4957f80-38f9-426e-86cb-37a07e3b7a1a · outbound

This paper cites BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T19:04:25.964826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:04:25.964826Z digest=sha256:ced75ba50a37bfeefa3d12f37b9a6370a8ad12200b3e55ef09cfc02ea856e756

Observation 841a835e-fbcf-427d-a7d2-e0249605e945 · outbound

This paper cites ACM Comput.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility ACM Comput

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:04:26.334240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:04:25.968242Z digest=sha256:ad2154ad7be2960c159d88d3430dd770eba5ea91a757eb48b9654ef039a06228

Observation e25778f0-993d-4db0-91f9-63d64667c6ec · outbound

This paper cites In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:04:26.324022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:04:25.971681Z digest=sha256:ce062bcccaf34b42efe8ff29bc619de32337c2eaf6248fab91ffbaeec97120e6

Observation 5538b631-824e-4882-be67-1604f0ec4d8a · outbound

This paper cites AMBROSIA: A Benchmark for Parsing Ambiguous Questions into Database Queries.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility AMBROSIA: A Benchmark for Parsing Ambiguous Questions into Database Queries

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T19:04:25.974940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:04:25.974940Z digest=sha256:46ee0a15dc9a4345dc9b8ea2e1e3b1df935297b5d61b2caedd9bc7438a9ebd06

Observation 641b435f-c8d6-4d26-aeaf-5bf8f6da3200 · outbound

This paper cites Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T19:04:25.978610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:04:25.978610Z digest=sha256:9711e9153f5d813a563eb8320df8bb4350a27fc8eb42add72b134bc1d8a743b1

Observation 41fb8a0d-318b-45b7-95ab-53e4389c1f5e · outbound

This paper cites Benchmarking TPU, GPU, and CPU Platforms for Deep Learning.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility Benchmarking TPU, GPU, and CPU Platforms for Deep Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T19:04:25.986001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:04:25.986001Z digest=sha256:d57551e3ec540d283f35bac7387accc305ae16520e052fd4a83697e8c20de9bf

Observation 6c0761f5-e342-418f-a39e-575223e02759 · outbound

This paper cites Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:04:26.312967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:04:25.989890Z digest=sha256:7ee3c15d9e0857ef2a6223c0b85c5affe83b4cfc7e3fd417108640fd5e692f39

Observation 8871a011-9d4f-4647-b182-379e83234441 · outbound

This paper cites Plot2Code: A Comprehensive Benchmark for Evaluating Multi-modal Large Language Models in Code Generation from Scientific Plots.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility Plot2Code: A Comprehensive Benchmark for Evaluating Multi-modal Large Language Models in Code Generation from Scientific Plots

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T19:04:25.993387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:04:25.993387Z digest=sha256:d429e0e0d840e5dc2be851619e005a3b01b7e92d04b5c97da88e501abcbdfa67

Observation b5a330f3-866f-42e8-8471-0b3810afb14f · outbound

This paper cites CRUXEval-X: A Benchmark for Multilingual Code Reasoning, Understanding and Execution.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility CRUXEval-X: A Benchmark for Multilingual Code Reasoning, Understanding and Execution

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T19:04:25.997038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:04:25.997038Z digest=sha256:c68db74010ed7969715ee0c252f69e925c455c21c6f3722a551d8f09f3723481

Observation c618f911-c028-42e7-98f4-592bb192423a · outbound

This paper cites HumanEvo: An Evolution-aware Benchmark for More Realistic Evaluation of Repository-level Code Generation.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility HumanEvo: An Evolution-aware Benchmark for More Realistic Evaluation of Repository-level Code Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T19:04:26.001569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:04:26.001569Z digest=sha256:d81620c7646e9ca5cd7a759d5cbae3d314321fd2dfc4cbb1378be6d731c4433c

Observation 4bbc60dd-80a8-42d0-82c1-d7fec200b412 · outbound

This paper cites 0 1000 2000 3000 0 10 20 30 40 50 60 70 80 90 100 110 120 130 140 150 160 170 180 190 Citations Figure 10: Citation Distribution of Benchmarks Coding Task.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility 0 1000 2000 3000 0 10 20 30 40 50 60 70 80 90 100 110 120 130 140 150 160 170 180 190 Citations Figure 10: Citation Distribution of Benchmarks Coding Task

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:04:26.301471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:04:26.009889Z digest=sha256:8608984678a51f777714a3b0365be4824393eab9f441c8cd8dbe67d566aed023

Observation 7144b574-bcae-4541-bb62-9410db5821aa · outbound

This paper cites For example, Hu- manEval (Chen et al., 2021a) and MBPP (Austin et al., 2021)), class-level (i.e., a class with mul- tiple function units of code.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility For example, Hu- manEval (Chen et al., 2021a) and MBPP (Austin et al., 2021)), class-level (i.e., a class with mul- tiple function units of code

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:04:26.290186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:04:26.013463Z digest=sha256:b2b3659cd741cc425b1a2762900f37b261cbda52cf9659f427180917f1453902

Observation e886fbfb-412f-468a-8bd7-cdf61fa6333e · outbound

This paper cites "" if len(dict.keys()) == 0: return False else: state =.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility "" if len(dict.keys()) == 0: return False else: state =

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:04:26.268643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:04:26.020641Z digest=sha256:18ebb24ed32f3a5a7a5beb5ede4c4abefaa556eaa1171883dc7f0c8d36598000

Observation e2308edd-b6d1-4238-a348-5085acc9db28 · outbound

This paper cites assert upper_ctr('PYthon') == 1.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility assert upper_ctr('PYthon') == 1

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:04:26.258180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:04:26.024098Z digest=sha256:ddcf2ed053f79811cd1103281a954d319a3904c8dcf5d711f8111b65a5544174

Observation a0345340-8e1f-40de-b16c-c684c63cdb84 · outbound

This paper cites CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis

Reference 294

Resolution
unresolved
no resolver link, observed 2026-08-10T19:04:25.953770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:04:25.953770Z digest=sha256:0bbee8466cb9ec4084f4762e42547c117a283d64e9ac8cfb4181950b12f4ee0a

Observation b87d9451-8b50-4d17-aa66-36f687992189 · outbound

This paper cites RES-Q: Evaluating Code-Editing Large Language Model Systems at the Repository Scale.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility RES-Q: Evaluating Code-Editing Large Language Model Systems at the Repository Scale

Reference 516

Resolution
metadata mismatch
local_arxiv, observed 2026-08-10T19:04:26.196889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:04:25.935489Z digest=sha256:bd4c4e3ddd8a8f5c17e0adfdea8c5ad351b09ac711ec86626d32cf48da957986

Observation 6e9b65b5-40ce-4d6e-bbf6-3c801369c065 · outbound

This paper cites The Impact of Reasoning Step Length on Large Language Models.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility The Impact of Reasoning Step Length on Large Language Models

Reference 1442

Resolution
unresolved
no resolver link, observed 2026-08-10T19:04:25.931274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:04:25.931274Z digest=sha256:aaeb05c8582209e9e9445dcf6bc9cf1bbda3da43fa46bab703e5dc8b83e177b0

Observation 24b41b94-ab75-4ba6-a038-b2b2c6862b19 · outbound

This paper cites Gorilla: Large Language Model Connected with Massive APIs.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility Gorilla: Large Language Model Connected with Massive APIs

Reference 1569

Resolution
unresolved
no resolver link, observed 2026-08-10T19:04:25.957389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:04:25.957389Z digest=sha256:212ea18b6906a0bd5d04e7ab02c6d2c1cc6df63b710cecde6fb461ac9bc9c965

Observation 9220f3cd-78d5-4726-b8e8-7a63997dcb24 · outbound

This paper cites VHDL-Eval: A Framework for Evaluating Large Language Models in VHDL Code Generation.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility VHDL-Eval: A Framework for Evaluating Large Language Models in VHDL Code Generation

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-10T19:04:25.982237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:04:25.982237Z digest=sha256:f2db916389dd7d709fcfe2f9ddb4e4a132b346774c1db95616a6f74af94b9cd5

Observation 8db7c4e9-ea1e-4336-94e0-7ecb703eb50f · outbound

This paper cites Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-10T19:04:26.005585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:04:26.005585Z digest=sha256:c85c3867b3a22431b39d771d70d9353526999d9ec163763abef102ea493f4cd3

Observation 7457eae2-0314-4877-87a8-8e5fa287585b · outbound

This paper cites Dianshu Liao, Shidong Pan, Xiaoyu Sun, Xiaoxue Ren, Qing Huang, Zhenchang Xing, Huan Jin, and Qinying Li.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility Dianshu Liao, Shidong Pan, Xiaoyu Sun, Xiaoxue Ren, Qing Huang, Zhenchang Xing, Huan Jin, and Qinying Li

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:04:26.345346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:04:25.943198Z digest=sha256:e1c06dc191748833dc764f5b6e13aed060cea91e11a57f623e9f41544eb93508

Observation c77e4a68-9820-48cf-837b-9db14ee61242 · outbound

This paper cites an unresolved cited work.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility Unresolved cited work

Reference 2019

Resolution
unresolved
raw_fallback, observed 2026-08-10T19:04:26.384535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:04:25.893263Z digest=sha256:c2cdf00a21c493311382d8f64d2bd2139e036d3c6ac7eb0dc36176310ea5d6c0

Observation bc2a4932-c16a-47b3-844d-063958e4e60f · outbound

This paper cites In Proceedings of the 28th International Confer- ence on Computational Linguistics, COLING 2020, Barcelona, Spain (Online), December 8-13, 2020 , pages 26–38.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility In Proceedings of the 28th International Confer- ence on Computational Linguistics, COLING 2020, Barcelona, Spain (Online), December 8-13, 2020 , pages 26–38

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:04:26.394682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:04:25.888664Z digest=sha256:ad2550ab7afb09baff52ec1d325e17cde32d47aec249c288b328cad29a87eb7d

Observation 5439fe26-5714-40ce-9e05-b0c5bb2310f4 · outbound

This paper cites measure the ability of these models to synthesize short Python programs from natural language descriptions.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility measure the ability of these models to synthesize short Python programs from natural language descriptions

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:04:26.279136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:04:26.016874Z digest=sha256:82886987bbe469026b149dd138c955098f654510a1aeec0fbaa048a9432fb38e

Observation 9e39f514-2992-4811-9d48-abf6cf58bbcd · outbound

This paper cites SySeVR: A Framework for Using Deep Learning to Detect Software Vulnerabilities.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility SySeVR: A Framework for Using Deep Learning to Detect Software Vulnerabilities

Reference 2022

Resolution
metadata mismatch
local_arxiv, observed 2026-08-10T19:04:26.181938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:04:25.939487Z digest=sha256:3c645559f60eb674981f0c8b2c1a69f92de0beb11c8bb209f6846e65e79c6f7a

Observation cc6295b4-01b7-4629-b9f5-3c910c840f3c · outbound

This paper cites In 45th IEEE/ACM International Conference on Software Engineering, ICSE 2023, Melbourne, Australia, May 14-20, 2023, pages 1430–.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility In 45th IEEE/ACM International Conference on Software Engineering, ICSE 2023, Melbourne, Australia, May 14-20, 2023, pages 1430–

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:04:26.356005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:04:25.927478Z digest=sha256:1dd8005cf75e18cf9c2f2fa662d566ac061f0efd34dc32ff4d3d78dcdc88e5c0

Observation b154828c-19f1-4b79-84bf-0bb7265d9243 · outbound

This paper cites Long Code Arena: a Set of Benchmarks for Long-Context Code Models.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility Long Code Arena: a Set of Benchmarks for Long-Context Code Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T19:04:25.910216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:04:25.910216Z digest=sha256:923a72789a5dd93dbd638bdbf9c55562e27522825088f6bf9dce0a6a27488d2f

Observation 43d11c34-c450-4efa-a791-c2a118d34726 · outbound

This paper cites Lakshya A.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility Lakshya A

Reference 5445

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:04:26.375676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:04:25.897709Z digest=sha256:491c616c4514b212c32ae46483e1bd75e1ae30dbbd60fac7d7c13d1cd34caa7b

Observation 35c6a7bb-d39d-49b3-a7d7-83b04d35f4e5 · outbound

This paper cites Ramakrishna Bairi, Atharv Sonwane, Aditya Kanade, Vageesh D.

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility Ramakrishna Bairi, Atharv Sonwane, Aditya Kanade, Vageesh D

Reference 8474

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:04:26.366160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:04:25.906189Z digest=sha256:df37c13387000c6852432e4be4a2497c85b8c61e983911d91c1de35d80a745e0

Pith citing papers

Observation beede259-c224-4d6f-b773-850b7dbae6e6 · inbound

Across Programming Language Silos: A Study on Cross-Lingual Retrieval-augmented Code Generation cites this paper.

Across Programming Language Silos: A Study on Cross-Lingual Retrieval-augmented Code Generation Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:18:46.014814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T11:52:44.990234Z digest=sha256:e39e73c4a67bb0a8bb33414dce690407f17c42a2afc3a6dde7604fe403edf432

Observation 3973a4fa-89d9-446d-a73e-f298d673980b · inbound

A Conceptual Framework for AI Capability Evaluations cites this paper.

A Conceptual Framework for AI Capability Evaluations Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:46.325101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:46.325101Z digest=sha256:2246099cdf47c204f3735658f0071d064c57f412ce81df8bea606e6075cfc282

Observation 58917dc5-a8db-4654-8d3a-2a65a00bfc8e · inbound

Establishing Best Practices for Building Rigorous Agentic Benchmarks cites this paper.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:20.949000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:20.949000Z digest=sha256:03a86bd1c869cfb5f876796173b2a2554d8039a74fc41a415289b66ecb768a61

Observation 6a12dde6-efe5-4dfd-a678-8026e8bc1c33 · inbound

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models cites this paper.

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-07T03:18:46.014814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:40:34.440305Z digest=sha256:e60cd5fdc959b4cdd86bf3af0a2ed3db0387919692d1ea79be71a4ce40ee7218

Observation fda34f98-829b-403e-be0e-948799f4c56c · inbound

Guidelines for Empirical Studies in Software Engineering involving Large Language Models cites this paper.

Guidelines for Empirical Studies in Software Engineering involving Large Language Models Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:18:46.014814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T22:02:36.307598Z digest=sha256:ade4f2f15b1b315fc49bad4af5ba741eac38a8bc8c8bd0af49c69f6949d7bcf3

Observation 65779e12-7cd7-49cc-bd71-1f333fd5523c · inbound

Guidelines for Empirical Studies in Software Engineering involving Large Language Models cites this paper.

Guidelines for Empirical Studies in Software Engineering involving Large Language Models Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:18:46.014814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T08:18:18.448122Z digest=sha256:cffb67b89c40654e2c60736d3804c771356a16c1a8355df9b70ee0d255f9a7c8

Observation 6631b1be-bbee-457c-9245-9f7bb63248dd · inbound

Knowledge-Graph-Driven Data Synthesis for Low-Resource Software Development: A HarmonyOS Case Study cites this paper.

Knowledge-Graph-Driven Data Synthesis for Low-Resource Software Development: A HarmonyOS Case Study Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-07T03:18:46.014814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:38:40.078229Z digest=sha256:a7ac75159beff4a57a6dd3930668e48835f14d7ff2b93d572e0c99f6b1f510c8

Observation 0061e5a0-8657-4136-9935-52d0364e35f4 · inbound

Edit, But Verify: An Empirical Audit of Instructed Code-Editing Benchmarks cites this paper.

Edit, But Verify: An Empirical Audit of Instructed Code-Editing Benchmarks Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-07T03:18:46.014814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T19:16:33.635516Z digest=sha256:99c561453dfbb29bb13db01b57353d3da5b42249fe653ec607f5ee802cbdc4bb

Observation 8513986c-d294-41b6-a051-13b72123b60d · inbound

Flaws in the LLM Automation Narrative cites this paper.

Flaws in the LLM Automation Narrative Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:18:46.014814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T10:52:36.252919Z digest=sha256:f841a3f39f8ad30a3fa03542dcb415d4072861773828a696bba77330a0572efd

Observation cc18cf3f-27f6-45f0-a7e8-ac73c421df8c · inbound

AlgoBench: Benchmarking Algorithmic Adaptation in Code Generation cites this paper.

AlgoBench: Benchmarking Algorithmic Adaptation in Code Generation Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-07T03:18:46.014814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-02T18:09:02.049387Z digest=sha256:e5ff11223682612905fd9fa25cda909014e352d9f3173b2d8b23c33e75e91a71

Observation eb8b0150-4c21-4608-81bf-219d2a213365 · inbound

To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing cites this paper.

To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T01:30:13.354497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T01:30:13.354497Z digest=sha256:a19598c8fbe7742325c069d6236ecb7d93c2989080f500a6ecf350500f574302