Pith. sign in

Paper Citation Record · LEDGER

Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2403.19114.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.19114 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:46:44.591883Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T08:34:27.439389Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e90f8d23-761a-4087-99c1-79c4c7dbde82 · inbound

Benchmark Data Contamination of Large Language Models: A Survey cites this paper.

Benchmark Data Contamination of Large Language Models: A Survey Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 166

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:10:41.129426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T23:10:40.420241Z digest=sha256:0135c7ec89658664d6abf22806583bcc65421e4bc77dde44bf9f7f028bd368fd

Observation ee0eebcf-5173-4998-9590-78a7c320bf4e · inbound

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities cites this paper.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.725599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.725599Z digest=sha256:9db6cdfef67e98a969b0b3bf041fd8c44eae6651f21bc9954bd75e8bc144798c

Observation 96d3e6c8-3e6f-4a8a-82ff-6496f9855eee · inbound

Unseen Horizons: Unveiling the Real Capability of LLM Code Generation Beyond the Familiar cites this paper.

Unseen Horizons: Unveiling the Real Capability of LLM Code Generation Beyond the Familiar Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T18:16:26.670461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:16:26.670461Z digest=sha256:bb55ba72b69b7259b87f2d99cef4ccd44ac9066b24d957a361be4459e32670a1

Observation 774795ee-b9cd-4c71-b77b-31099f857353 · inbound

CodeMorph: Mitigating Data Leakage in Large Language Model Assessment cites this paper.

CodeMorph: Mitigating Data Leakage in Large Language Model Assessment Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:37.247818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:37.247818Z digest=sha256:265aabcc503312d9b8b0f229e3466feb87cf340eed420692cd1d51682bbb8ef3

Observation d2961752-69f7-4042-b9f0-395cb5a083b5 · inbound

Combining TSL and LLM to Automate REST API Testing: A Comparative Study cites this paper.

Combining TSL and LLM to Automate REST API Testing: A Comparative Study Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:38.548821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:26:38.548821Z digest=sha256:2ab0634f578df76a8ce28213fc2160d513fe83d65714548a4231b02e17b16ba8

Observation 774f1bd7-85c3-42f8-b5e7-8a24ba23a3bd · inbound

Sustainable Code Generation Using Large Language Models: A Systematic Literature Review cites this paper.

Sustainable Code Generation Using Large Language Models: A Systematic Literature Review Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 144

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:50:16.506793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T18:49:01.097179Z digest=sha256:156fc1c41b190e518c551c8cbd37ade2faadd8ecce430b759100418512cf027e

Observation 4effb991-0933-4ee7-a21f-50f16cdbac37 · inbound

SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents cites this paper.

SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:21:27.205216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T06:13:32.434201Z digest=sha256:93f6edb91ff51a9d05901fa66372d8a3ff5b142fbf4f2ed06679243095ced43c

Observation 6824477f-af6f-4a10-ad7f-2b49a9c628a8 · inbound

FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale cites this paper.

FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:03:29.013236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T02:02:25.597640Z digest=sha256:1981c549d7768a1ebe06c7a4440f3141094acd4da8086e82dc942f86fda8801f

Observation b7683286-aac9-48aa-a80a-85bb01283e22 · inbound

Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation cites this paper.

Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 56

Resolution
malformed identifier
arxiv_id, observed 2026-06-30T08:34:27.441561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T08:30:37.016856Z digest=sha256:a2c562ab59ebe6e651492c1a2063f627e6bc66c4afd71144c81ffe266a1ae677

Observation 6945d7be-c82c-4611-9500-12bf084d04fc · inbound

Attention to Detail: Evaluating Energy, Performance, and Accuracy Trade-offs Across vLLM Configurations cites this paper.

Attention to Detail: Evaluating Energy, Performance, and Accuracy Trade-offs Across vLLM Configurations Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-13T04:56:20.002453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T04:56:20.002453Z digest=sha256:3511611b73deae3707d50e25609ee05d423fd487413c2d6df85757b2443d91ad

Observation 8440dad5-8fea-4ce1-9b80-024b1dfbed09 · inbound

Attention to Detail: Evaluating Energy, Performance, and Accuracy Trade-offs Across vLLM Configurations cites this paper.

Attention to Detail: Evaluating Energy, Performance, and Accuracy Trade-offs Across vLLM Configurations Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T07:43:20.935793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:43:20.935793Z digest=sha256:8ea35250288f987da56b9d6f7506fd36c8e9cae107a39e5588c06ca0796e80d5

Observation 1aea9a71-11e1-49af-a2d3-bd5bec753b54 · inbound

Memorization Diagnostics for Code LLMs Should be Scale-Aware cites this paper.

Memorization Diagnostics for Code LLMs Should be Scale-Aware Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T23:46:44.591883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:46:44.591883Z digest=sha256:b71ed700e93f4cc416ba8247a9c18b9d166fe0be72ff8cbe2a15183b65b38260